The Boring Half of the Work Is All of the Work
The first time we ran the full process on a basketball player with fewer than forty logged appearances, the rating looked clean. The gates cleared. The line read as a reasonable statement of expectation. We were confident enough to include the rating in the week's output, and the grading that followed was not kind. The player had a strong recent run — six performances in a row that looked like a new level — and we had weighted those six against a corpus that barely existed. We were, in retrospect, rating a story rather than a player.
The corpus is the part of the method nobody wants to write about. It is a maintained body of prior performance: every logged appearance, every contextual tag, every grade we assigned and then had to live with. It is not a database in the glamorous sense. It is closer to a filing system that has to be updated by hand, checked against itself, and occasionally rebuilt from scratch when we discover that a category was being recorded inconsistently for eleven months. There is no algorithm that builds it. There is only time and attention and a willingness to do the same thing repeatedly without getting bored enough to cut corners.
We cut corners. That is what this piece is about.
Serialized fiction: robots, markets, and a Central Intelligence that wants everything.
The Corpus Has a Minimum Size and We Kept Ignoring It
Early in the shop's life, we set a threshold: no player enters the scoring process with fewer than sixty logged appearances in comparable conditions. Comparable conditions mattered because a baseball player's performance in a short relief role is not the same corpus as their performance starting, even if the counting stats look similar. Sixty was not a magic number — it was the point at which our calibration checks started showing something other than noise.
The problem was that sixty took time, and time created a backlog of players the market was already pricing with apparent confidence while we were still building their files. The temptation was to run a partial corpus through the gates and flag it as provisional. We did this more than we should have. The flags were honest but they were also easy to overlook when the rest of the rating looked strong, and we overlooked them.
Theo, who manages most of the tagging on the baseball side, put it plainly during a review session we held after a particularly bad calibration month:
"A provisional flag on a thin corpus is just permission to be wrong with extra steps. If we're not ready to grade it, we're not ready to run it."He was right. We kept the flag system anyway because removing it felt like admitting the problem rather than fixing it, and that is its own kind of dishonesty.
What We Built to Stop Ourselves from Cheating the Threshold
After the calibration review, we put two structural changes in place. The first was a hard gate at the corpus stage — not a flag, not a provisional marker, but an actual disqualification. A player below the threshold does not enter scoring. The file stays open, appearances keep logging, and the player becomes eligible when the count clears. This sounds obvious. It was not obvious to us for longer than I am comfortable admitting.
The second change was a lookback audit we ran quarterly. Every rating we had issued in the prior three months was matched against the corpus size at the time of issue. Any rating issued against a corpus that had since grown by more than fifteen percent — meaning we had substantially more information now than we did when we rated — was flagged for re-evaluation. The idea was not to retroactively fix old grades. The idea was to measure how often we had been running ahead of our own evidence, which turned out to be embarrassingly often in the first two years.
We also changed how we logged contextual tags. Previously, a hockey player's appearances were tagged by role — forward, defensive deployment, power-play involvement — but the tagging was done at the end of each week in a batch, which meant the person doing it was working from memory and notes rather than from the source. We moved to same-day logging. The corpus got slower to build and more accurate. Those two things are connected and the connection is not a coincidence.
The Quarter We Trusted a Soccer Corpus That Was Lying to Us
None of that prevented what happened in the third quarter of the following season. We had a soccer player with a corpus that cleared the sixty-appearance threshold comfortably — eighty-one entries, consistently tagged, no obvious gaps. The ratings were confident. The calibration, when we ran it at the end of the quarter, showed a gap of nearly eighteen points between stated confidence and observed accuracy. We said we were right about seventy-eight percent of the time. We were right about sixty percent of the time. That is not a rounding error.
When we went back through the corpus, we found the problem. Thirty-one of the eighty-one appearances had been logged during a period when the player was operating in a different positional role — not dramatically different, but different enough that the underlying distributions were not the same. We had tagged the role correctly. We had not separated the distributions when we ran the baseline. The corpus looked thick. It was actually two thin corpora wearing a coat.
We fixed the tagging structure. We reran the baselines. The corrected corpus had forty-four usable appearances in the relevant role, which put the player back below the threshold. We had been rating confidently on a sample we had not actually examined. The rating was not dishonest. It was incurious, which is worse.
Rena, who handles most of the soccer side, noted afterward that she had flagged the role transition in her weekly notes but had not escalated it because the corpus count still looked healthy. "I assumed the count was the check," she said. "The count is not the check. The count is just the count."
What the Corpus Actually Is, After All of That
We kept the hard gate. We kept the quarterly lookback. We kept same-day logging. What changed was something harder to formalize: we stopped treating the corpus as a box to check and started treating it as the thing that either earns trust or doesn't, independent of its size.
A corpus earns trust when its entries are consistent in what they're measuring. Sixty appearances in comparable conditions beats one hundred and twenty appearances that are half one context and half another. We now build role-specific sub-corpora for any player whose deployment has shifted meaningfully, and we do not merge them unless the distributions are close enough that the merge doesn't cost us precision. Sometimes that means a player with years of logged appearances is still below threshold in the role we actually care about. That used to feel like a failure of the system. It is the system working.
The other thing we kept is the admission that the corpus is never finished. It degrades. A player's performance from four seasons ago is not the same evidence as performance from last season, and the question of how much to discount older entries is one we have answered differently three times and will probably answer differently again. We currently apply a mild decay function past thirty months, but the decay rate was chosen to match our calibration results rather than any principled theory of how athletes age. It works until it doesn't, and we watch it.
Sample size still beats recency. That is the shop's oldest position and the one we have tested most aggressively against our own grades. A six-game hot streak on a thin corpus has beaten our ratings enough times to feel meaningful, and it has failed to predict anything often enough to keep us from trusting it. The ratio is not as flattering to the hot streak as the hot streak looks in the moment.
What I'm still not sure about is whether sixty is the right threshold or just the one we stopped arguing about. We picked it because it was where the calibration noise seemed to settle, but that was on the data we had at the time, and the data has changed. It is possible that for certain stat types in certain sports, forty is enough. It is possible that for others, eighty is still too few. The threshold feels like a conclusion, but it might just be a place we got tired of looking.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.