A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

A Column That Passed Every Check and Said Nothing

For about four months last year, we ran a grading column that produced clean results every single week. Hit rate above the threshold we had set. No missing entries. Graded inside the window, logged on time, reviewed at the Tuesday meeting without incident. It was, by every measure we had built into the process, a success. Remi called it "the quietest column we've ever run," and meant it as a compliment.

It was not a success. It was a column that had learned to pass its own checks, which is a different thing entirely and considerably worse. The accuracy looked fine because the column was rating candidates where the outcomes were easy to predict — not because the method was working, but because we had, without noticing, drifted toward a corner of the corpus where we were least likely to be embarrassed. The grades came back clean. The calibration was a disaster.

We caught it late, which is the honest version of how this story goes. The tell was not a bad week — it was a suspiciously good month, followed by a review in which we could not explain, even to ourselves, what the column had actually measured.

Get Smarter Fantasy Sports Insights, Rankings, and Lineup Tools

Access fantasy rankings, projections, lineup optimizers, waiver advice, and expert analysis across football, baseball, basketball, and more.

Learn more

How a column passes every check and still measures nothing

The column had started as a focused effort on basketball player performance in secondary statistical categories — not the obvious ones, but the kind of outputs a corpus can accumulate quietly over a long season. We had a reasonable sample. The gates were applied correctly. The shrinkage was in place. On paper, the setup was sound.

The problem was in what we were grading against. We had defined accuracy as whether the rated direction matched the observed outcome — a binary, pass or fail, logged in a column that fed the weekly summary. What we had not done was track how confident we were in each rating, or whether that confidence was distributed evenly across the candidates we were taking on. We were measuring whether we were right. We were not measuring whether we deserved to be right.

This distinction sounds obvious when I write it down. It was not obvious in the room. We had a process that produced a number, the number looked acceptable, and the Tuesday meeting moved on. The column that we forgot to check for six weeks at least had the excuse of neglect. This one we were actively watching, and we still missed it.

What we built to fix grading, and how it fooled us anyway

When we first designed the grading structure, we were careful about the window. Fixed length, no extensions, no retroactive adjustments. We had learned from an earlier episode — a window we kept moving to protect a rating that was simply wrong — and the fixed window was the direct response to that. It worked, in the sense that we stopped moving the goalposts.

What we had not built was a column that tracked the difficulty of each rating. We had a hit-rate column. We had a sample-size column. We did not have anything that recorded whether a given candidate was hard to call or easy, whether the corpus behind it was deep or thin, or whether our stated confidence was proportionate to what the evidence actually supported. The column was measuring outcomes. It was not measuring whether our confidence was earned.

Remi flagged this first, in a way that I initially dismissed. "Every hit this month came from candidates where we already knew the answer," she said during a review. I told her that was fine — that knowing the answer before grading was the point. She let it go. She was right and I was wrong, and it took another six weeks to see why.

"The column isn't grading the method. It's grading the candidates we were confident enough to run through the method. That's a different population entirely."
— Remi

She had identified, in one sentence, the selection problem we had built into the process without realising it. We were not randomly sampling from the candidate pool. We were implicitly filtering toward the cases where we felt comfortable, and then grading ourselves on those cases. The accuracy looked good because we had, quietly and without any formal decision, stopped taking on the hard ones.

The four months of clean results that told us we were well-calibrated when we weren't

The cost was not a single bad miss. It was four months of data that we now cannot use for calibration purposes, because the population the column graded was not representative of the candidates the method is supposed to handle. We ran the retrospective and found that the difficult candidates — thin corpus, contested lines, categories with high variance — had been trickling out of the column since around week three. Not removed deliberately. Just quietly deprioritised at the point where we chose which candidates to take through the full process.

This is the version of the problem that does not feel like a mistake while it is happening. It feels like good judgment. You look at a candidate with a thin corpus and you think: not enough data, skip it. You look at one with a contested line and you think: too much noise, skip it. Each individual decision is defensible. The aggregate effect is that you have built a grading column that only sees the cases it was already going to get right.

We had also, separately, added a confidence-logging field to the grading sheet earlier in the year — a field that was supposed to record our stated confidence before each outcome was known. That field had been left blank for most of the four months in question. We had built the instrument and then not used it, which is a particular kind of failure. It is the kind we wrote about when we added a column to feel thorough rather than to actually be thorough. We had done the same thing again, in a different direction.

The hit rate for those four months was, depending on how you count, somewhere between twelve and eighteen points above our historical baseline. At the time we noted this with quiet satisfaction. In retrospect, a hit rate that far above baseline should have been the alarm, not the reward. When a process starts outperforming its own history by that margin, the correct question is not "what are we doing right?" It is "what have we stopped counting?"

What the column looks like now, after we rebuilt the grading criteria

We kept the fixed window. That part was never the problem. We kept the binary outcome column — right or wrong, no partial credit, no narrative explanation that softens the miss. Those two elements are load-bearing and we did not touch them.

What we added was a difficulty tier, logged at the point of candidate selection, before the outcome is known. Three levels, defined by corpus depth and line clarity. Every candidate gets a tier. The grading summary now reports hit rate by tier, not in aggregate. A clean aggregate hit rate with all the hard cases sitting in tier three and never being graded is now visible as exactly that.

We also reinstated the confidence-logging field and made it mandatory — not optional, not a nice-to-have. If the field is blank, the entry does not count toward the weekly summary. This is a small mechanical fix with a disproportionate effect: it forces the person running the column to commit to a stated confidence before the outcome is known, which is the only version of calibration that means anything. The number that turns out to measure nothing is almost always a number that was never stress-tested against a stated prior.

The hit rate dropped immediately when we reintroduced the hard candidates. It dropped by more than we expected, which told us the four clean months had been doing more damage to our self-assessment than we had accounted for. We are now, by our own grading, performing below the threshold we had set for the column. That is, perversely, the most honest result we have produced in about six months.

The thing I keep returning to is that the column was not broken in any way we had thought to check for. Every gate cleared. Every field populated. Every deadline met. The failure was upstream of all of it — in the quiet, unrecorded decisions about which candidates were worth the effort. I don't know how you build a gate for that. I'm not sure you can.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top