Shrinkage We Applied Evenly and Shouldn't Have
Shrinkage is supposed to stop you believing yourself. The idea is simple enough: take your estimate, pull it back toward the average, and accept that your confidence is probably inflated. We built a version of that rule early, applied it to every sport on the same terms, and did not seriously question the uniformity for longer than I am comfortable admitting. The rule felt principled. It was principled. It was also wrong in a way that took two sports pulling in opposite directions before we noticed.
The specific failure was this: we were shrinking basketball estimates and tennis estimates by the same coefficient. Basketball, where a roster of twelve means one player's absence redistributes minutes across several others and the base rate of any individual stat is genuinely unstable night to night. Tennis, where there is no roster, no redistribution, and a player's serve percentage across five hundred service games is about as stable a number as this shop ever handles. Pulling both toward their respective means at the same rate is not a neutral act. It is importing basketball's noise assumptions into a sport that does not have them.
We caught it in a calibration review — the kind of quarterly check where you compare stated confidence against observed accuracy and find out whether you were right for the right reasons. The tennis estimates were systematically under-confident. The basketball estimates were fine. We had been shrinking away a real signal in tennis for months, and the uniform rule was the reason.
Securely manage keys for 60+ AI providers in one encrypted vault instead of juggling them across apps.
Why the Same Coefficient Means Different Things in Different Sports
Shrinkage works by acknowledging that extreme estimates are more likely to be the product of noise than of genuine signal. The further your estimate sits from the mean, the harder you pull it back. The strength of that pull should be proportional to how noisy the underlying data is — how much of what you observed was real and how much was variance that will not repeat.
Basketball is noisy in a specific, structural way. Minutes fluctuate. Matchups rotate. A player who logged thirty-four minutes on Tuesday may log twenty-two on Thursday because a foul situation changed or a blowout shortened the game. The sample that feeds any individual estimate contains a lot of that kind of variance, and shrinking hard toward the base rate is a reasonable response to it.
Tennis does not have that structure. A player serves every point they play. There is no coach deciding to pull them after fifteen minutes, no teammate absorbing their workload, no lineup card. The variance that exists is real — form fluctuates, surfaces matter, opponent quality shifts — but it is smaller in magnitude and more predictable in source than basketball variance. A corpus of four hundred matches for a player who has been on tour for six years is telling you something reliable. Pulling that estimate aggressively toward the tour average is not humility. It is discarding information you earned.
Remi put it plainly when we were working through the calibration numbers:
"We built the shrinkage rule after a bad basketball stretch. We were overcorrecting for basketball's noise and then we just... kept the correction and handed it to every other sport."That was accurate. The rule had a history we had mostly forgotten, and the history was sport-specific.
What We Tried When We Noticed the Tennis Estimates Were Soft
The first response was not to change the shrinkage coefficient. It was to look for something wrong with the tennis corpus — to ask whether the sample was smaller than we thought, or whether there was a gate failing to screen out low-information matches. This is a reasonable instinct, and it was wrong. The corpus was fine. Most candidates had cleared the gates correctly; the ones that survived had the sample depth we require. The under-confidence was not a corpus problem. It was a shrinkage problem.
Once we accepted that, we ran a sport-by-sport decomposition of the calibration gap — how far off stated confidence was from observed accuracy, broken out by sport. Basketball: within acceptable range. Hockey: slightly over-confident, which we had seen before and attributed to the difficulty of reading ice time. Soccer: roughly neutral. Tennis: under-confident by a margin that was not noise. The tennis number had been sitting there in the quarterly reports, and we had been reading it as sampling variation. It was not sampling variation. It was a systematic pull in the wrong direction.
The fix we tried first was blunt: we raised the tennis shrinkage floor — the minimum estimate we would allow before shrinkage pulled it further — rather than changing the coefficient itself. The logic was that raising the floor would stop the most extreme downward pulls without requiring us to recalibrate the whole rule. It helped at the margins and did not solve the problem. The coefficient was still doing work it should not have been doing.
What the Uniform Rule Actually Cost, and Where We Were Wrong About It
The honest accounting is that we published a quarter of tennis ratings that were softer than the evidence warranted. We did not publish them as soft — we published them at whatever confidence level the model assigned, which was lower than it should have been because the shrinkage had already done its damage upstream. The ratings were not wrong in direction. They were wrong in confidence, and shrinking confidence on principle is only virtuous when the principle is correctly calibrated to the sport.
There is a version of this error that is forgivable: you apply a rule uniformly because you do not yet know it should not be uniform, you find out, you fix it. That is how method development works and we are not embarrassed by the sequence. What we are embarrassed by is that the signal was visible in the calibration reports for two quarters before we read it correctly. We saw the tennis under-confidence number, filed it as noise, and moved on. That is not a corpus failure or a gate failure. That is a reading failure, and it sits with us.
The floor adjustment we tried first cost us an additional month of sub-optimal ratings before we admitted it was a partial fix. I think we tried it because changing a floor feels smaller than changing a coefficient — less like admitting the rule was wrong and more like adjusting a parameter. The distinction is cosmetic. We were adjusting the rule either way; we just chose the version that felt less like a reversal.
What We Changed and What the Sport-Specific Coefficient Actually Looks Like Now
We now set shrinkage coefficients per sport, derived from each sport's observed within-season variance relative to its between-season variance. Sports where a player's performance in week three predicts their performance in week eleven poorly get more aggressive shrinkage. Sports where that prediction holds up get less. Tennis, as expected, sits at the low end. Basketball sits near the high end. Hockey is complicated by ice time in a way we are still working through.
The practical result is that tennis estimates move less from their corpus-derived starting point before they are published. They are pulled toward the mean, because all estimates should be, but the pull is proportional to what the sport's own variance history justifies rather than to a single coefficient that was calibrated against basketball noise. The shrinkage floor is now also sport-specific, which is the change we should have made when we tried the floor adjustment the first time.
What we kept from the original rule is the principle: shrinkage is not optional, and confident estimates get pulled toward the base rate on purpose. The shop's oldest mistake is believing a strong corpus reading too completely, and the rule exists to prevent that. We kept the rule. We stopped pretending the rule could be sport-agnostic when the sports themselves are not.
Remi's note in the revision document read:
"The coefficient was always a guess. We just forgot we were guessing."That is a fair summary of how most of our rules start and why the calibration reviews exist.
The question we have not fully resolved is whether sport-specific coefficients create a new version of the original problem — a set of assumptions we derived from one era of data in each sport and will apply too long into the next one. A tennis player's serve percentage may be stable across five hundred matches until it isn't, and the coefficient we set to reflect that stability will be slow to notice when the stability breaks down. We traded one uniform assumption for several sport-specific ones, which is an improvement, but it is not the same thing as having no assumptions.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.