A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Column We Forgot to Check

The column was called conf_stated. It sat in the grading sheet between the outcome column and the rolling hit-rate column, and for six weeks in the back half of last season's basketball stretch, nobody opened it. Not because we forgot the sheet existed. We ran the sheet every week. We updated outcomes, recalculated the hit rate, noted the embarrassing ones. We just never scrolled right.

The hit rate looked fine. That was the problem. It hovered in a range that felt honest — not suspiciously high, not so low that anyone would ask questions. So we called the grading current and moved on. What we were not checking was whether the confidence we had stated at the time of each rating matched the accuracy we were actually observing. That is a different question entirely, and we had built the column specifically to answer it. We had then, apparently, decided the column could answer it alone in the dark.

When Priya finally pulled the full sheet to prep a calibration review, she sent a message that said only: "conf_stated. When did you last look at this." It was not a question. We had been overstating confidence by a consistent margin for six weeks without knowing it, and the hit rate had been quietly covering for us the whole time.

Gear Up for Your Next Outdoor Adventure

Shop hunting and outdoor gear including waders, boots, backpacks, apparel and accessories built for the field and everyday outdoor use.

Learn more

How a column you built yourself becomes invisible

The conf_stated column records the confidence level we attached to a rating at the moment we issued it — before any outcome was known. The idea is simple: if we said 70% confident across a set of ratings, we should be right about 70% of the time on that set. If we are right 85% of the time, we were underconfident. If we are right 55% of the time, we were overconfident. Neither direction is neutral. Overconfidence is the one that erodes the whole premise of the shop, because the shop's premise is that it knows how much it knows.

We had added the column after an earlier stretch where our stated confidence and our observed accuracy had drifted badly in the other direction — we had been underselling ratings that turned out to be well-supported, which sounds modest but is its own kind of miscalibration. That piece of history is worth reading if you want the full arc of how that column came to exist in the first place, because the reason we built it matters to understanding why we stopped checking it.

The reason, as best I can reconstruct it: the column had been stable for a long time before this stretch. Month after month, the gap between stated confidence and observed accuracy sat within a range we considered acceptable. Stable things stop attracting attention. The sheet became a ritual rather than an instrument, and rituals do not require you to read every part of them.

What we thought we were doing for those six weeks

The grading workflow during that stretch was not negligent on its face. Outcomes were logged within the fixed window — that discipline held. The hit-rate column was current. We had a brief weekly review where someone would scan for obvious anomalies and flag anything that needed a second look at the corpus. Two ratings were pulled back for re-examination during the period. The process looked like it was running.

What we were actually doing was grading accuracy while ignoring calibration. Those are related but they are not the same thing. Accuracy tells you how often you were right. Calibration tells you whether your confidence in being right was itself reliable. A shop that is right 68% of the time and said 68% is in better shape than a shop that is right 72% of the time and said 90%. We knew this. We wrote about it. We then ran a six-week stretch behaving as though the hit rate was the only number that mattered.

"The hit rate is what you show someone who doesn't know the shop. Calibration is what you show yourself. You stopped showing yourself."
— Priya

She was right, and the phrasing stung specifically because it was accurate rather than harsh. We had been performing the grading process for an imagined external audience rather than running it for the shop's own use. That is a drift that happens slowly and is very hard to catch from inside.

Six weeks of overconfidence, quantified badly

When Priya pulled the full sheet, the picture was consistent and uncomfortable. Across the six-week window, our stated confidence averaged roughly 11 percentage points above our observed accuracy. The gap was not random noise. It tracked a specific pattern: ratings issued on the back of recent strong performance in the corpus were being assigned confidence levels that the full historical sample did not support. We had been, in effect, letting a warm stretch in the data talk us into certainty that the long-run record had not earned.

This is the failure mode we have written about before under a different name. Sample size beats recency — it is the shop's oldest and most-tested belief — and yet here was six weeks of grading data showing that recency had been quietly inflating our stated confidence without anyone catching it. The corpus was not wrong. The line had been signaling something the corpus underweighted, and rather than treating that as a reason for more caution, we had apparently treated it as confirmation.

The harder cost was to the grading record itself. We keep every rating, including the embarrassing ones, and we do not quietly delete misses. That discipline is the shop's whole claim to credibility on this subject. But a grading record that shows accurate outcomes and miscalibrated confidence is a partial record, and we had let it be partial for six weeks. The outcomes were correct. The confidence column was not. Calling that "fine" because the hit rate looked acceptable was a form of selective reading we had not previously caught ourselves doing.

I want to be precise about what the error was not: it was not a data error, and it was not a methodology error. The method for assigning confidence had not changed. The column existed and was being populated. The failure was attentional — we stopped treating one part of the sheet as load-bearing, and it turned out to be load-bearing.

What the sheet looks like now, and what we are still not sure about

The immediate fix was structural. The weekly review now requires a named person to state the current calibration gap aloud — not in writing, aloud, in the room — before the session closes. Writing it down had clearly not been sufficient. The spoken version creates a moment where someone has to have actually read the number rather than scrolled past it. It is a low-tech solution and we are not proud of needing it.

We also tightened the confidence-assignment step at the rating stage. Any rating that draws on a corpus window of fewer than forty observations now has a hard ceiling on stated confidence, regardless of what the recent performance looks like. The ceiling is not generous. This was already nominally policy; it had been applied inconsistently during the stretch in question, which is how the recency inflation crept in. Consistent application of an existing rule is less satisfying to announce than a new rule, but it is what the problem actually called for.

What we kept, unchanged, is the fixed grading window. There was a brief conversation about whether extending the window would smooth out some of the calibration noise, and we decided against it for the same reasons we have always decided against it — the window has to be fixed in advance and it has to stay fixed, or the grading process becomes something you can adjust to protect a rating you liked. The six-week problem was not a window problem. Changing the window to address it would have been the wrong repair on the wrong part.

The question we are still sitting with is subtler. The calibration gap during those six weeks was consistent in direction — always overconfident, never under. That kind of systematic drift suggests something structural rather than random. We think we know what it was: the recency inflation, the under-application of the confidence ceiling. But a consistent directional gap is the kind of thing that makes you wonder whether there is a second cause you have not found yet.

A calibration gap you cannot explain in one sentence is probably not explained yet. We have a sentence. We are not entirely sure it is the whole sentence.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top