A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

When More Data Made the Rating Smaller

We had a basketball player in the corpus who looked, at 40 games, like a genuinely unusual performer. His per-possession output in a particular counting category sat well above the positional baseline. The confidence interval was wide, as it always is at that sample size, but the point estimate was high enough that the entry cleared our internal threshold and started influencing ratings. We were pleased with ourselves, which is usually the first sign something is about to go wrong.

By the time the entry reached 120 games, the estimate had moved — not corrected for an injury, not adjusted for a role change, not revised because we found a data error. The underlying numbers were fine. We had simply added more games, and the more games we added, the closer the estimate settled toward the positional mean. The player was not declining. The corpus was converging. Those are different things, and for a few weeks we were not fully sure which one we were watching.

The experience was useful in the way that mildly embarrassing things usually are. It forced a cleaner articulation of what our shrinkage step is actually supposed to do, and it surfaced a quiet assumption we had been carrying about what "enough data" means — an assumption that turned out to be wrong in a specific and reproducible direction.

Learn to Analyze Data Like a Front Office

A free online course in data analytics with Python: statistics, visualization and finding the signal in the numbers.

Learn more

Why the Entry Looked Bigger Than It Was at 40 Games

The short version is that 40 games is enough to clear our minimum gate but not enough to distinguish a genuinely elevated performer from a player who ran warm for a third of a season. We knew this in the abstract. The corpus doctrine says it plainly: sample size beats recency. What we had not fully internalized was how aggressively a small sample can mislead a point estimate upward specifically, rather than in either direction randomly.

The mechanism is not mysterious. Early in a corpus entry, the player's actual performance dominates the estimate because there is not much prior to pull against. Our shrinkage step was applying a positional prior, but at 40 games the weight on the prior was modest — we had calibrated the blending schedule to give observed data room to speak. That calibration was correct in principle. The problem was that we were also, informally, reading the high point estimate as confirmation that the player was special, rather than reading it as an estimate that had not yet been tested by volume.

Remi noticed it first. She had been tracking a cluster of entries that all came in high early and then drifted, and she flagged the pattern in a note that was more diplomatic than it needed to be.

"The ones that look best at forty games are not the ones that look best at a hundred and forty. I don't think that's survivorship. I think we're reading early noise as signal and then being surprised when the noise goes away."

She was right. We had seen this pattern before in entries we eventually stopped trusting without being able to say exactly why — and in retrospect, the mechanism was the same one. Early inflation, slow convergence, and a reluctance on our part to revise downward because downward revision feels like admitting the player is worse than we thought, rather than admitting the estimate was always uncertain.

How We Tried to Stabilize the Entry as Games Accumulated

The first thing we did was check for a structural break — a point in the timeline where something about the player's situation changed and the later data was measuring a different thing than the earlier data. Role change, new teammates, different usage patterns. We found nothing convincing. The player's context was stable. The counting category we were tracking had not been redefined or recollected. The data feed was clean.

The second thing we did was look at whether our blending schedule was weighting the prior too lightly at intermediate sample sizes. This is a real failure mode — we had written about it in the context of building a corpus around the wrong version of a stat, where the prior was essentially decorative because the observed data overwhelmed it too quickly. We ran the entry through a version of the model with a heavier prior at the 40-to-80 game range and compared the trajectories. The heavier prior would have produced a lower, more stable estimate from the start. It would also have suppressed a handful of entries that genuinely were elevated and stayed elevated. We did not want to fix one problem by creating another.

The third thing — and this is the one that actually changed how we work — was to look at the variance of the estimate across the first 40 games, not just the point estimate. The variance was high. We had been reporting the point estimate prominently and the interval quietly, and the interval was telling a different story than the point estimate. An interval that wide, centered on a number that high, should have read as "we don't know yet" rather than "this player is good." We had been reading it wrong.

What the Inflation Actually Cost the Ratings That Depended on It

The honest accounting: for roughly six weeks, the player's corpus entry was pulling his composite rating above where it should have been. The ratings that used that entry as an input were correspondingly elevated. We did not catch this in real time. We caught it during a calibration review when we noticed that a cluster of ratings involving this player had been systematically overconfident — we had said 70% confidence more often than the outcomes warranted, and the gap was not random noise.

The calibration gap was not large enough to be catastrophic, but it was consistent enough to be diagnostic. When a gap is consistent in one direction across a cluster of related entries, it is almost never a coincidence. It is a shared upstream assumption that is wrong. In this case the shared assumption was that a high point estimate at 40 games was informative rather than provisional.

What made it worse was that we had documentation saying the opposite. The shop's position on sample size is not ambiguous — it is the oldest belief we hold and the one we have tested most. The failure was not ignorance of the principle. It was a quiet exception we had made for this particular entry because the early numbers were flattering, and flattering numbers are harder to be skeptical of than bad ones. This is, I suspect, the same psychology that keeps entries in rotation long past the point where they are doing useful work — we are more willing to keep something that looks good than to interrogate it.

We also lost about three weeks of clean calibration data for that player's category because we could not separate "the rating was wrong because the corpus entry was inflated" from "the rating was wrong for some other reason." When the upstream estimate is noisy, the grading window downstream becomes harder to interpret. That cost was small in absolute terms and annoying in principle.

What the Converged Entry Looked Like and What We Changed Around It

At 120 games, the entry settled at roughly 8% above the positional baseline — elevated, but modestly. The confidence interval had narrowed considerably. The estimate was, in the language we use internally, believable: a number we could defend by pointing to volume rather than by pointing to a run of good games. The player was a slightly above-average performer in this category. That is a useful thing to know. It is not as exciting as what the 40-game entry was suggesting, but it is more likely to be true next season than the earlier number was.

The structural change we made was to the way we display intermediate estimates. We already reported confidence intervals; we added a second indicator that flags any entry below 80 games as provisional in the internal dashboard. The flag does not change how the entry is used in calculations — that would require a more significant redesign — but it changes how we read the output. A provisional entry with a high point estimate now reads as a question rather than a finding. That is a small change with, so far, a measurable effect on how often we reach for early-stage entries to support a conclusion we already want to reach.

We also revisited the blending schedule. We did not change the weights — the heavier-prior version had too many side effects — but we documented the specific range (40 to 80 games) where the current schedule is most likely to produce inflated point estimates, and we added that range to the gates review. The gate does not disqualify entries in that range. It flags them for a second read before they influence a composite rating. The gate we almost added — a hard cutoff at 80 games — felt clean on paper and would have removed a real class of useful early-career data. We did not add it. Sometimes the right gate is a warning, not a wall, and knowing the difference between those two things is most of the work.

The entry is still in the corpus. It is rated, graded, and calibrated like everything else. What I am still not sure about is whether the problem was the blending schedule, the display, or something simpler — that we wanted the player to be exceptional and the early data gave us permission to believe it, and no process catches that particular failure mode reliably.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top