Shrinking Our Own Confidence on Principle
For a long time we tracked hit rate and called it good. The rating either cleared the threshold or it didn't, the outcome either matched the direction or it didn't, and at the end of a grading window we counted the column that said yes and divided by the total. This is the most natural thing in the world to do, and it is also how you miss the actual problem for months at a stretch.
The problem is not whether you are right. The problem is whether your confidence in being right matches how often you actually are. A shop that says it is 80% sure and lands 80% of the time is doing something real. A shop that says it is 80% sure and lands 62% of the time is doing something else entirely — it is telling a story about its own certainty that the outcomes keep quietly contradicting. We were doing the second thing. We just hadn't built the machinery to see it.
The piece that follows is about what we did when we finally looked. It is not a story about fixing ourselves. We are not fixed. It is a story about shrinkage — the deliberate, mechanical act of pulling your own confident estimates back toward the base rate before they go anywhere near a grade — and what it felt like to apply that to confidence levels themselves rather than just to the underlying numbers.
Create short links, track clicks, and understand your audience. Privacy-friendly by design. No cookies, no tracking pixels, just the stats you need.
The Gap Between "80% Sure" and What 80% Sure Actually Looks Like
Calibration, stated simply: if you make a hundred claims at 70% confidence and you're right on 70 of them, you are calibrated. If you're right on 53 of them, you have a gap of 17 percentage points that you owe an explanation for. The gap is not a rounding error. It is a structural claim about how well your confidence tracks reality.
We started measuring ours seriously after we graded ourselves generously for months without noticing. The generous grading was partly definitional — we had been allowing ourselves to count near-misses as hits under certain conditions — but the deeper issue was that our stated confidence levels had never been empirically derived. They came from feel. A rating that cleared the gates cleanly got called high-confidence. A rating that scraped through got called moderate. Nobody had ever checked whether those labels predicted anything.
When we ran the first proper calibration pass, the results were not catastrophic but they were consistent in one direction: we overclaimed. Our high-confidence ratings landed at a rate meaningfully below what "high confidence" implies. Our moderate ratings were actually closer to accurate, probably because the hedging built into the label was doing work that the underlying model wasn't. Nadia, who runs most of our grading infrastructure, put it plainly in the debrief.
"The moderate bucket is honest because it sounds modest. The high-confidence bucket is broken because it sounds like we know something."
She was right. And the fix was not to recalibrate our model. The fix was to stop trusting the model's own self-assessment and apply shrinkage to it on principle.
Applying Shrinkage to Confidence, Not Just to Estimates
Shrinkage, in the way we use it, means pulling an estimate toward the base rate in proportion to how much you distrust your sample. You have a player with a short recent history who looks exceptional — shrinkage says that exceptional-looking number is probably partly noise, so you move it toward the population average. The more uncertain your data, the harder you pull. This is not pessimism. It is arithmetic.
We had been applying this to our player estimates for years. The idea of applying it to our own stated confidence levels was newer and felt stranger, because confidence is supposed to be the thing doing the correcting, not the thing being corrected. But the calibration data made the case. If our high-confidence bucket was systematically overclaiming, then the right response was to treat "high confidence" as an estimate subject to the same shrinkage logic as anything else.
The implementation was mechanical. We set a prior based on our historical calibration gap by bucket — how far off each confidence tier had run over the previous grading windows. Then we moved every stated confidence level toward the base rate by that gap before it was recorded. A rating that would have been logged at 82% confidence came out at something closer to 71% after adjustment. We did this uniformly, before seeing any new outcomes, which matters: writing the threshold down before you see the data is the only way to prevent the adjustment from becoming a post-hoc story.
The first few weeks felt like we had broken something. Ratings that had previously looked authoritative now looked appropriately tentative. Several of us found this uncomfortable in a way that was hard to articulate. The ratings hadn't changed. The outcomes hadn't changed. Only the honesty of the confidence label had changed, and somehow that was the part that felt like a loss.
What We Got Wrong About What Shrinkage Was Fixing
We thought shrinkage on confidence would improve our calibration scores quickly. It did not. The calibration gap narrowed, but it narrowed slowly — over several months rather than several weeks — and in the meantime we had introduced a new problem we hadn't anticipated.
By pulling all our high-confidence ratings toward the moderate range, we had compressed the distribution. We now had very few ratings at the extremes. This sounds like appropriate humility, and in one sense it is, but it also meant that when a rating genuinely deserved high confidence — when the corpus was deep, the gates were clean, and the historical match rate was strong — we were still shrinking it by the same prior that had been derived from our worst overclaiming. We were penalizing good data for the sins of bad data.
This is the part that a week of being confidently wrong across multiple sports had originally obscured: the overclaiming problem was not uniform. It was concentrated in specific conditions — sparse corpora, contested availability, stat types that vary wildly by opponent. Applying a flat shrinkage prior across all conditions meant we were now underclaiming in exactly the situations where we had earned some confidence, and the calibration data eventually showed that too.
We had solved one asymmetry by introducing another. The lesson we took — and I use "lesson" loosely because we are still working through it — is that shrinkage priors need to be conditioned on the same variables that drive the original overclaiming. A flat penalty is better than no penalty, but it is not a finished answer.
What the Shop Kept, and What It Still Argues About
We kept the shrinkage. We did not keep the flat prior. The current version conditions the pull on corpus depth, gate passage score, and stat-type volatility — the three variables that had predicted overclaiming most reliably in the historical data. Ratings built on shallow history get pulled harder. Ratings built on deep, clean corpora get pulled less. The distribution is no longer compressed at the center.
We also kept the calibration reporting. Every grading window now produces a calibration chart by confidence tier, and that chart is the first thing reviewed before anything else is discussed. This sounds like a small procedural change. It has turned out to be the most argumentative fifteen minutes of any review cycle, because the chart has a way of surfacing disagreements about method that politeness had previously buried. Marcus, who handles the basketball and hockey corpus work, described it as "the meeting where we find out what we actually believe versus what we said we believed." That is an accurate description.
The thing we still argue about is whether any of this changes the underlying ratings or only changes how we talk about them. The shrinkage prior adjusts the confidence label. It does not adjust the estimate itself — the number that describes what we think a player is likely to do. Those two things are related but not the same, and there is a live internal dispute about whether a rating with a shrunken confidence label is a different product or just an honestly labeled version of the same one. I have a position on this. So does everyone else, and the positions are not the same.
What is not in dispute is that a good stretch of results is the most dangerous time to skip the calibration check. The months when everything is landing are exactly the months when overclaiming quietly accelerates, because the outcomes aren't there to push back. The shrinkage prior runs regardless of recent results. That is the part we are most confident about, which means it is probably the part we should be watching most carefully.
There is a version of this work where you get the priors right, the calibration closes, and the confidence labels finally mean what they say. We have not arrived there. What I keep returning to is whether calibration is a destination at all, or whether it is more like a pressure — something you apply continuously against your own tendency to believe yourself, knowing the tendency will keep returning. The chart does not answer that. It just tells you where you stood last month.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.