Shrinkage Floors Set Differently Per Sport
For the first two years, we used a single shrinkage floor across every sport we tracked. The logic seemed clean: pull every estimate toward the base rate by a fixed amount, cap the most confident ratings, and let the calibration step tell you whether you went too far. One number, one rule, consistent application. We were proud of the tidiness. We were also wrong in a way that took embarrassingly long to surface.
The problem was not that the floor was set too high or too low. The problem was that "the base rate" means something structurally different in a sport where the top performers play eighty-two nights a year versus one where they play four rounds over a fortnight. In basketball, a large corpus exists almost by accident — the schedule generates it. In tennis, you are always negotiating with scarcity, and a floor calibrated for volume will sand down the signal from a player whose sample is genuinely thin for reasons that have nothing to do with uncertainty. We were treating a measurement problem as if it were a philosophy problem.
What followed was about eighteen months of uncomfortable re-examination. We did not rebuild the method from scratch — the structure held. But we had to accept that the cap-and-shrink step was doing different work in different sports, and that pretending otherwise was a form of false precision we had not noticed because the errors were spread across sports rather than concentrated in one.
Securely manage keys for 60+ AI providers in one encrypted vault instead of juggling them across apps.
Why a single floor punished tennis and flattered basketball
Shrinkage, as we practice it, is a correction for overconfidence. When an estimate sits far from the base rate, we pull it back. The further it sits, the harder we pull. The floor is the point below which we will not pull — the minimum distance we allow between an estimate and the mean, because some real signal deserves to survive the correction.
In basketball, a player who has appeared in sixty games this season has a corpus that can absorb a fairly aggressive floor. If the estimate is still far from the base rate after sixty observations, that distance is probably real. The floor can be set low — meaning we pull hard — without destroying genuine signal, because the sample is large enough to have earned its deviation.
In tennis, sixty appearances might represent three full seasons. A player deep in a draw, on a surface they rarely play, against opponents outside their usual range — the corpus is thin not because the player is unknown but because the sport is sparse. Setting the same floor here means either pulling too hard and flattening real distinctions, or setting a floor so high that it stops doing any work at all. We tried both. Neither was right.
Hockey sat somewhere in the middle and gave us a different headache: the corpus was large enough, but the variance in a single game was high enough that the base rate itself was unstable. A floor calibrated to basketball's cleaner distributions was systematically too tight. We kept getting estimates that looked precise and turned out to be measuring noise we had mistaken for signal. Nadia put it plainly during one of our review sessions:
"We built a floor for a sport where the mean is stable. Then we applied it to a sport where the mean moves four points depending on which goaltender dressed. Those are not the same problem."
She was right, and we had been slow to see it.
How we tried to set sport-specific floors without overfitting them
The obvious fix — just tune a separate floor for each sport until the calibration looks good — is also the most dangerous one. If you set the floor by looking at results, you are not calibrating a method; you are fitting a parameter to historical data and calling it a principle. We have written about the temptation to set grading windows after seeing results, and sport-specific shrinkage floors carry exactly the same risk.
What we tried instead was to anchor each floor to a structural property of the sport rather than to its outcomes. The properties we settled on were three: typical season length (as a proxy for how much corpus a player can accumulate in a given window), within-game variance for the stat type in question, and the degree to which a single contextual factor — opponent, surface, lineup — could move the base rate substantially. Each property was scored on a simple scale, and the floor was derived from the composite score rather than from the hit rate.
This meant the floor for a tennis surface-split stat was set higher — meaning we pulled less aggressively — than the floor for a basketball volume stat, not because tennis had better calibration history but because the structural argument for restraint was stronger. We were trying to make the decision before we saw the answer, which is the only version of the decision we trust.
We also introduced a corpus-size modifier that interacted with the floor. Below a threshold of appearances — we settled on thirty as the minimum for any estimate to be scored at all, a number we arrived at by arguing about it rather than by testing it — the floor was automatically raised regardless of sport. The gate logic we had developed for availability checks turned out to be useful here as a model: a hard disqualification below the threshold, a graduated adjustment above it.
What the sport-specific floors got wrong in the first cycle
The structural anchoring worked better than the single universal floor. It did not work as well as we had hoped, and the failure was instructive.
The within-game variance score turned out to be harder to measure cleanly than we expected. We were estimating it from the same corpus we were using to build the ratings, which introduced a circularity we did not catch until the first full grading cycle came back. The variance estimate was being inflated by outlier performances — exactly the performances that shrinkage is supposed to discount — and so the floor was being set higher than it should have been in high-variance sports. We were protecting signal that was actually noise, because the noise had inflated our estimate of how noisy the sport was.
In practical terms, this meant our hockey estimates were less shrunk than they should have been. The ratings looked confident. The calibration gap — the distance between stated confidence and observed accuracy — was wider than in any other sport we tracked that cycle. We had built a correction for overconfidence and then introduced a new source of overconfidence inside the correction itself.
There is a version of this problem that is almost funny: the mechanism designed to stop you believing yourself too much can be gamed by a bad input into believing itself too much. The variance estimate is upstream of the floor, which is upstream of the estimate, and if you are not careful the whole chain is self-referential. Marcus spent a week untangling the hockey corpus from that cycle and concluded that roughly a third of the floor adjustments had been set in the wrong direction. We published the miss. It is still in the grading record.
What the revised floors look like now, and what they still cannot do
After the hockey correction, we rebuilt the variance estimate from a held-out portion of the corpus — data we had deliberately not used in building the ratings — rather than from the full set. This introduced a lag: the floor is always calibrated to a window that ends before the current one begins. That felt like a loss at first. It turned out to be a feature. The floor cannot be contaminated by the very estimates it is supposed to discipline.
The floors now sit at four distinct levels across the five sports we cover, with basketball at the most aggressive end and tennis at the most restrained. Baseball and soccer share a middle tier, for different structural reasons — baseball because the corpus is large but the within-game variance for individual stats is genuinely high, soccer because the base rates for most individual stats are low enough that small absolute errors produce large relative ones. Hockey sits in its own tier above baseball and soccer, still the sport that punishes overconfident floors most visibly.
What the floors still cannot do is account for the cases where a corpus entry clears every gate and still carries a structural lie — a player whose historical numbers are real but whose current role has changed in a way the corpus has not yet absorbed. The floor will not save you from a corpus that is technically complete and practically stale. That is a different problem, and we have not solved it.
The floors also do not interact with the line-reading step in any formal way. The two mechanisms run in parallel, which means it is possible for a shrinkage floor to preserve an estimate that the line has already revised past. We have noticed this more in tennis than anywhere else, where a single match result can move beliefs faster than a corpus update cycle can follow. Whether the floor should be responsive to line movement is a question we have not answered to our satisfaction.
The thing we are least sure about is whether the structural anchoring is actually more principled than outcome-tuning, or whether it just feels more principled because the reasoning happens before we see the results. The calibration record does not clearly distinguish between them yet. We have a preference, and we act on it, but I am not certain the preference is anything more than a prior we have not tested hard enough.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.