PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Accuracy and Calibration Are Different Things

Two shops can be right the same proportion of the time and one of them can be worth listening to while the other is not.

The difference is whether the stated confidence means anything. It took us most of two years to understand that this is the actual product, and that accuracy on its own is close to worthless as a claim about quality.

Where test results become engineering insights

Analytics, observability, and AI-driven insight from test runs.

Read iTestResults

Our First Scorecard Only Had One Column

For a long time we tracked one thing: what proportion of our ratings turned out correct. It went up, we were pleased. It went down, we investigated. That was the entire feedback loop.

The problem is that a single proportion cannot tell you whether your confidence is honest. A shop that rates everything at moderate confidence and is right two-thirds of the time is behaving well. A shop that rates everything at very high confidence and is right two-thirds of the time is behaving badly, and the two are indistinguishable on a scorecard with one column.

We were the second shop. Not egregiously — our confidence ran perhaps a little ahead of our accuracy, consistently, in a way that never looked alarming in any individual month. But the direction never changed, which is what makes it a bias rather than noise.

The reason it persisted is that our accuracy was genuinely fine. Every time somebody asked whether the method worked, we could point at a defensible number and the conversation would end there. The number was true and it was answering a question nobody should have been satisfied with.

Bucketing Our Own Confidence and Grading Each Bucket

The fix is standard and we should have had it from the beginning. Group every rating by the confidence we assigned it. Within each group, measure how often we were actually right. Compare the two.

Done that way, the picture is immediately legible. Our moderate-confidence ratings were close to honest. Our most confident ratings were the worst calibrated by a wide margin — we were right substantially less often than we had claimed we would be, precisely where we were most emphatic.

That pattern turns out to be common and it has an obvious mechanism once you see it. High confidence requires several components to agree, and components that share an input agree more often than independent ones would. So the most confident ratings were the ones where our redundancies were doing the most work.

“Your ninety-fives are behaving like your seventies,” Priya said, the first time she drew the curve. “Everything else is roughly fine. It's only the top that's lying.”

Where the Correction Had to Come From

The honest fix was not to improve the confident ratings. It was to stop issuing them.

We compressed the top of our confidence range — the highest levels we had been using are simply no longer available to the scoring, because we could not produce a rating that deserved them. That felt like a downgrade and in a narrow sense it was one. Our headline confidence numbers got worse. Anyone comparing us to a shop willing to say ninety-five would find us less impressive.

It also cost us an argument we had been winning internally for a year. The most confident ratings were the ones people found most interesting to discuss, and removing that tier removed a lot of the pleasure from the weekly review. Meetings got duller. That is a real cost and I would not pretend otherwise.

And it did not fix the underlying cause, which is component dependence. It bounded the damage. The dependence is still there and we have made only partial progress on it.

The Curve Is the Scorecard Now

We do not report a hit rate as a headline any more. The scorecard is the comparison between stated confidence and observed accuracy, by bucket, and the single number we care about is how far apart those two run at the top end.

The rule that came with it: a confidence level we cannot support is not available to be used. If a bucket is persistently overconfident, the bucket goes, rather than being adjusted and kept.

What I would say to anybody building something similar is that accuracy is the number you will be asked for and calibration is the number that tells you whether you are any good. We spent two years reporting the first one honestly and it functioned as a kind of self-deception anyway.

The curve is close to the diagonal now except at the very top, where it still sags slightly. I do not expect that to fully go away, and I have stopped treating it as a bug to be closed.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top