A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

When the Corpus Grew and the Ratings Got Worse

For most of the shop's first year, the corpus was thin. We knew it was thin. Every rating that came out of it carried an implicit asterisk — the sample was small, the confidence intervals were wide, and we said so in the grading notes. Then we spent the better part of four months acquiring two additional seasons of historical performance data across basketball and baseball, and we fed all of it in. The corpus nearly tripled in size. We expected the ratings to sharpen. They got measurably worse.

The calibration drop was not enormous. It was not the kind of thing that announces itself in a single bad week. It showed up gradually in the grading logs: stated confidence at 72%, observed accuracy at 58%. Then 71% against 57%. Then, in one particularly grim stretch, 74% against 54%. We had been running closer to parity for months. Something in the expansion had introduced a systematic lean, and we did not know what it was for longer than I would like to admit.

What followed was about six weeks of pulling the corpus apart and putting it back together. Most of what we found was not dramatic. It was the kind of structural problem that only becomes visible when the data is large enough to make the signal noisy — which is its own small irony, because the whole point of building a larger corpus was to reduce noise.

Know What Your Business Can Actually Spend

See the cash truly available after bills, payroll, taxes, and reserves before making your next move.

Learn more

How a Bigger Corpus Quietly Broke the Ratings

The corpus we built in year one had a hidden property we had not noticed: it was accidentally coherent. Because it was small, it drew from a narrow window of seasons, which meant the statistical environment it described was relatively stable. Scoring rates in the basketball data, for instance, had not shifted dramatically within that window. The base rates we were shrinking toward were plausible approximations of what the players were actually doing.

When we added two older seasons, we added data from a different offensive environment. The per-game averages were lower. The pace was different. The players who had long careers bridging the old data and the new data appeared in the corpus twice — once as a younger, lower-output version of themselves, and once as the player we were actually trying to rate. The corpus was treating those as two independent observations of the same underlying talent, and averaging them together. It was not. It was averaging two different players who happened to share a name.

This is a version of the problem we have written about before — a corpus entry that keeps shifting is often a sign that the underlying records are not describing a stable entity. We had seen that symptom in a single player's file. We had not seen it at scale, distributed invisibly across hundreds of entries, until the calibration scores started drifting.

What We Tried First: Era Flags and Weighted Recency

The first fix we tried was era-flagging. We split the corpus into two periods — call them the old environment and the current one — and applied a recency weight that discounted the older seasons. The idea was straightforward: keep the historical data for sample-size purposes, but tell the model to trust it less when projecting into the current environment.

Nadia built the weighting scheme over about a week. The decay function was conservative — we did not want to throw away the older data entirely, because the whole premise of the shop is that sample size beats a hot streak, and gutting two seasons of history to fix a calibration gap felt like it violated the spirit of the thing. So we split the difference: older seasons contributed to the corpus at roughly 40% of the weight of recent ones.

"The problem with a decay function is that it gives you a knob to turn, and once you have a knob you will turn it until the backtest looks good. That's not fixing the problem. That's tuning to the test." — Nadia

She was right, and we turned the knob anyway. The calibration improved slightly — stated 71%, observed 63% over the next three weeks — but it did not recover to where we had been. We had treated a structural problem as a weighting problem, and the weighting fix was masking the real issue rather than resolving it.

The real issue, which took us another two weeks to isolate, was that the era split was not the only source of contamination. We also had a stat-version problem layered underneath it. Some of the historical data used an older definition of a counting stat — one that had since been revised — and we had ingested it without checking. This is a failure mode we have documented in detail elsewhere, and the embarrassing part is that we had already documented it. We knew this could happen. We did not check carefully enough when the new data came in.

What the Expansion Actually Cost Us

The direct cost was six weeks of degraded output. Every rating produced during that window carried a calibration gap we were not fully aware of for the first two weeks and could not fully explain for another four. We graded all of it. None of it was deleted. It sits in the logs as a long, uncomfortable stripe of overconfident ratings, and it will affect our rolling calibration numbers for the rest of the year.

The subtler cost was to our confidence in expansion itself. We had treated corpus growth as unambiguously good — more data, more signal, better ratings. That assumption was wrong in a specific and instructive way. More data is better when it describes the same underlying process. When it describes a related but distinct process — an older offensive environment, a revised stat definition, a player at a different career stage — it can introduce more noise than it removes. The corpus is not a bucket you fill. It is a model of a process, and adding data that does not fit the process makes the model worse.

Marcus, who handles most of the grading review, put it plainly when we did the post-mortem: "We celebrated the acquisition and skipped the audit." That is accurate. We spent four months getting the data and about four hours checking whether it was compatible with what we already had. The ratio was wrong.

There is also a quieter lesson embedded in the calibration numbers that I keep returning to. The ratings did not get dramatically wrong — they got confidently wrong. The stated confidence went up slightly as the corpus grew, because larger samples produce tighter estimates. But the observed accuracy went down, because the larger sample was partly measuring something different. Accuracy and calibration are not the same thing, and this was a clean demonstration of why: we had a period where we were more confident and less right at the same time.

What We Kept After Rebuilding the Expansion Logic

We kept the additional data, but we rebuilt the ingestion logic around it. Every new source now goes through a compatibility check before it touches the main corpus — era alignment, stat-definition version, career-phase tagging for players whose output changed significantly over the period in question. The check is not automated in any sophisticated way. It is mostly a structured manual review with a short checklist. It takes about a day per data source. We should have been doing it from the start.

We also kept the era-weighting, but we recalibrated it against held-out data rather than against the backtest. Nadia's warning about knob-turning was correct, and the version of the decay function we are running now was set before we looked at the calibration improvement, not after. Whether that discipline will hold the next time we are staring at a gap and a knob is nearby — I genuinely do not know.

The gates got a new entry. Before any candidate is scored, we now check whether the corpus records for that player span more than one statistical era, and if so, whether the era-weighting has been applied to their file explicitly. It is a small gate. It catches maybe one candidate in fifteen. But it is the kind of check that would have caught the problem we spent six weeks diagnosing, and a gate that closes after the damage is done is still a gate worth closing.

What we did not do is cap corpus size or set a ceiling on how far back we look. The instinct to do that was strong during the worst weeks of the calibration drop — just throw out everything older than three seasons and rebuild from clean data. We resisted it, partly because it felt like overcorrecting, and partly because the shop's oldest belief is that sample size beats recency. Abandoning that principle because a particular expansion went badly would be the wrong kind of lesson to take.

The question I have not resolved is whether the compatibility check is actually doing what we think it is, or whether it is just a more elaborate version of turning the knob until the numbers look acceptable. We built the check to catch the specific failure we experienced. The next failure will probably be a different shape, and it is not obvious that a checklist designed for the last problem will see the next one coming.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top