PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Why a Proven Performer Outranks a Hot Streak

If you took everything else away and left us one rule, it would be this one: a long, ordinary record is worth more than a short, excellent one.

It has survived four rebuilds of the scoring, two changes of sport coverage, and every argument anybody has brought against it. It is also wrong in two specific situations, and it took us a long time to find them.

Test data engineering for modern systems

Generation, validation, and management of test data at scale.

Read iTestData

The Streak Always Looks Like the Better Candidate

Recent form is the most emotionally available evidence there is. It is what everybody is discussing, it is what the coverage is about, and it feels like information about *now* rather than information about the past.

Structurally, though, a streak is a short sample selected for being unusual. That is the whole definition. You noticed it because it stands out, which means you have selected on the outcome, which means the thing you are measuring is contaminated by how you found it.

A long ordinary record has none of that. Nobody selected it. It is just what somebody does, observed many times, and the estimate it supports is far less exciting and far more likely to hold next week.

Weighting History Against Recency, and Getting It Backwards Twice

Our first attempt weighted recent appearances more heavily, because that is what everybody does and it felt obviously correct. It made the estimates worse. Measurably, over a full season.

Our second attempt removed recency entirely. That also made things worse, though less, and for a reason we did not anticipate: some changes in a player's situation are real and permanent, and a model with no recency at all cannot notice them.

Where we ended up is unglamorous. History dominates. Recency is allowed to move the estimate only when it persists long enough to stop being a streak — which, by construction, means we are always slow to notice a genuine change. We accept that cost explicitly, because the alternative cost us more.

“We are going to be late on every real improvement,” Marcus said, when we settled it. “I'd rather be late on the real ones than early on the fake ones, and there are a lot more fake ones.”

What finally convinced everybody was not an argument, it was a back-test we did not want to run. We took two full prior seasons, scored every candidate both ways, and graded both sets against what actually happened. The recency-weighted version was worse in every sport we covered, and it was worst precisely where it felt most convincing — on players in the middle of a visible run of form. That result is now the only reason anybody in the shop believes this, and I mention it because the reasoning on its own had failed to persuade us for about a year.

The Two Cases Where the Rule Is Simply Wrong

It is wrong when a player's role has changed structurally. If somebody's usage genuinely doubles, their long ordinary record is now a record of a different job, and weighting it heavily is not conservatism — it is measuring the wrong thing carefully. We were slow on this for a whole season and produced a run of confidently poor estimates, all of them defended by the rule.

It is also wrong for anybody early enough in their career that the long record does not exist to be weighted. We spent a while treating those players as unrateable, which is defensible, and a while treating them as ordinary, which is what the shrinkage does. Neither is satisfying. There is genuine information in a short record and our method throws nearly all of it away.

I do not have a fix for the second one. We have looked at it three times and each attempt has made something else worse, so it sits in the notebook as a known limitation rather than a solved problem.

The Rule, With Its Exception Written Next to It

The rule stands, and the role-change exception now sits directly beside it in the doctrine, with a definition of what counts as a role change that we wrote before we needed it.

What we did not do is soften the rule to accommodate the exception. That was the tempting move — blend the two, weight recency a little more, split the difference. We tried that. It produced a method that was mildly bad at both jobs instead of good at one and explicitly blind in a named situation.

Being explicitly blind somewhere is a much more comfortable position than being vaguely mediocre everywhere. It is also easier to write down, which means the next person to look at it can see the hole.

Somebody will eventually show me a version of this that handles both cases without giving up anything. I have been waiting for that for a while and it has not arrived, which is weak evidence that it is hard.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top