PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Season We Trusted a Nine-Game Sample

Nine appearances. That was the entire body of evidence, and for about three weeks it produced the highest-scoring candidate in the system.

Nothing was broken. Every component did what it was built to do. That is what made it worth writing up — the failure was in the design, and the design was ours.

Where test results become engineering insights

Analytics, observability, and AI-driven insight from test runs.

Read iTestResults

A Short Record With No Variance In It

The player had been available for nine outings and had been remarkably consistent across all nine. Not spectacular. Consistent.

Consistency is exactly what our scoring rewarded, because consistency is what makes a performance predictable, and predictability was the whole point. A player who does roughly the same thing every time is a player you can say something about. So the score came back very high, and it came back high for a reason that was, narrowly, correct.

The trouble is that nine observations of anything look consistent. You need a good deal of history before inconsistency has room to show up. What we had measured was not this player's reliability — it was the near-certainty that a short record contains no surprises yet.

There is a second thing a short record hides, and it is the one that actually caught us. Variance is not just absent from nine observations — it is unmeasurable from them. Our scoring did not distinguish between a player whose spread we had measured and found narrow, and a player whose spread we had not yet had the opportunity to measure at all. Both arrived at the scoring stage as a single number with no attached uncertainty, and once that number is in hand there is nothing left in the pipeline to remind anybody how thin the evidence behind it was.

Pulling Confident Estimates Back Toward the Middle

The fix is old and boring and we should have had it from the start: pull every estimate toward the population's typical behaviour, and pull harder when the evidence behind it is thin.

A player with four seasons of history barely moves. A player with nine appearances moves a long way, because we genuinely do not know much about them and the honest estimate is close to "probably ordinary". It is not a penalty for being new. It is a correction for how little we have observed.

The effect on that particular candidate was immediate — they dropped out of the top of the list entirely and settled somewhere unremarkable, which is where a player we know almost nothing about belongs.

“It's not that the nine games were lying,” Dana said. “It's that nine games can only tell you a small thing, and we were reading a large thing off them.”

Three Weeks of Being Confidently Impressed

For three weeks that candidate was the thing we talked about. It shaped what we looked at, what we thought the model was good at, and — this is the part that stings — what we thought we had discovered about finding undervalued players early.

We had discovered nothing. We had found a short record and a scoring function that could not tell the difference between a reliable player and an unobserved one. And because the candidate kept performing consistently for another week or so, we got positive feedback for the error, which is the worst possible outcome. A mistake that immediately fails is cheap. A mistake that works for a fortnight teaches you the wrong lesson thoroughly.

When it did break, it broke unremarkably: the player had an ordinary bad outing, then another, and the record started to look like everybody else's. Nothing dramatic. The evidence simply arrived and it was not what we had assumed.

Why New Players Are Now Deliberately Boring to Us

Shrinkage applies to everything now, always, with no exception for a candidate somebody finds interesting. That last clause is doing real work: the exceptions were always requested for the candidates we were most excited about, which is precisely when the correction matters most.

We also stopped treating a high score on a thin record as a signal worth investigating. It is not an interesting anomaly. It is the expected output of a short history, and if it appears we now read it as a reminder that the shrinkage is doing its job rather than as something to look into.

Sample size beats recency. We had that written down before any of this happened. Having it written down turns out to be different from believing it.

I still know who that player was. They have had a perfectly ordinary career since, which is the least dramatic possible ending and the one the corrected estimate predicted.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top