A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Column We Added to Feel Thorough

Sometime in the middle of a slow baseball stretch, we added a column to our grading sheet that nobody had asked for. It was called "context-adjusted outcome" and it lived between the raw grade and the final calibration tally. The idea was that a rating deserved partial credit if the player had performed directionally well but an unusual game situation had pulled the counting stat in the wrong direction. A pitcher who threw seven strong innings but inherited a blown save from a reliever. A basketball player whose assists were suppressed because his team went to isolation sets down the stretch. We wanted to capture the signal beneath the noise.

The column felt serious. It looked thorough. For about four months, we filled it in conscientiously, and it made our grading meetings longer and more satisfying in the way that meetings become satisfying when everyone has a lot to say. What we did not do — and this is the part I find genuinely difficult to explain — was ask whether the column was connected to anything. It fed no downstream calculation. It influenced no shrinkage parameter. It was, in the most precise sense available, decorative.

We kept it anyway, because removing it felt like admitting we had wasted the time we spent filling it in. That is the trap, and I don't think it is unique to analytics shops.

Understand Why Modern Life Feels So Hard

Explore clear explanations for the money, work, technology, health, and relationship problems we face every day.

Learn more

Why a Column That Measures Nothing Still Feels Like Measurement

The honest version of the problem is this: grading is uncomfortable. When a rating misses, the miss sits in the sheet and you look at it. The raw grade does not care why the miss happened — it records that your stated confidence was not matched by observed accuracy, and it moves on. That is correct behavior, and it is also somewhat brutal if you have a plausible explanation for the miss.

The context-adjusted column was, in retrospect, a pressure valve. It let us write something next to the bad grade — something that gestured at nuance — without actually changing the grade. We told ourselves it was documentation. What it was, more accurately, was a place to put our feelings about the miss so they didn't contaminate the number.

The problem with that arrangement is that it contaminated the number anyway, just more slowly. When you spend fifteen minutes in a grading meeting discussing why a miss deserves partial credit, you are training yourself to remember the miss differently. The raw grade says the rating failed. The column says it almost didn't. Over time, the column's version of events is the one that sticks. This is not a small problem for a shop whose entire premise is that it publishes its own error rate honestly.

Nadia put it plainly during one of the later reviews:

"We built a column that argues with the grade. That's not documentation. That's an appeal process with no one to appeal to."
She was right. The column was a second opinion we were giving ourselves, and we were always the most sympathetic reviewer available.

What We Tried When We Realized the Column Was Doing Nothing

The first instinct was to make the column do something — to wire it into the calibration calculation in a way that gave partial credit a formal weight. If a rating missed because of a provably unusual game situation, maybe the calibration penalty should be halved. We spent about two weeks designing that system, which required us to define "provably unusual" in a way that couldn't be gamed by the person doing the grading.

We could not do it. Every definition we wrote had an escape hatch. "Unusual game situation" became "any situation the rater found surprising," which is another way of saying "any situation where the rating was wrong." The partial-credit system was, structurally, a miss-forgiveness system, and a miss-forgiveness system is precisely what a calibration process cannot contain. The grading window has to be fixed in advance — and for the same reason, the grading criteria have to be fixed in advance too. Deciding after the fact what counts as an excusable miss is the same error wearing different clothes.

The second instinct was to keep the column but make it purely qualitative — a note field, not a grade modifier. We tried this for another six weeks. What happened was that the notes became longer and more elaborate, and the grading meetings became longer still, and we were no closer to understanding whether our confidence levels were accurate. We had built an excellent record of our own rationalizations.

What Four Months of Thorough-Looking Work Actually Cost Us

The direct cost was time, and the time was not trivial. Filling in the column, discussing it, defending the entries, and occasionally arguing about whether a specific game situation qualified as unusual — conservatively, this added forty minutes to each weekly grading session. Over four months, that is a meaningful fraction of the hours we had available for calibration work.

The indirect cost was harder to quantify and probably larger. The column gave us a story about our misses that was more flattering than the raw data supported. When we eventually ran the numbers — comparing our stated confidence levels against observed accuracy for the period when the column was active — the gap was wider than in the preceding period. We had gotten less calibrated while doing more grading work. I want to be careful not to claim the column caused that gap; there were other variables, including a stretch where we were running the same model across sports it wasn't built for and not catching the drift. But the column did not help, and the mechanism by which it hurt seems clear enough.

The thing we got unambiguously wrong was the assumption that more documentation produces more honesty. It can. It can also produce more sophisticated dishonesty, which is worse because it is harder to detect. A single raw grade is hard to argue with. A raw grade plus a context-adjusted entry plus a note field plus a fifteen-minute discussion is very easy to argue with, and the argument almost always favors the rater.

This is adjacent to a mistake we had made before — resetting a grading window after a miss embarrassed us — where the mechanism was different but the impulse was identical: finding a procedurally legitimate way to make the record look better than it was.

What Stayed in the Sheet and What We Removed

We removed the context-adjusted column entirely. Not archived, not moved to a supplementary tab — deleted. The grading sheet now contains the rating, the outcome, the raw grade, and the calibration delta. That is all. If someone wants to write a note about why a miss happened, they can do it in a separate document that is explicitly labeled as post-hoc analysis and is not attached to the grade record.

The separation matters. Post-hoc analysis is useful — understanding why a rating missed is how you improve the model. But that analysis has to live somewhere that cannot retroactively influence the grade, because the moment it can, you have built an incentive to write favorable post-hoc analysis. The shop's whole claim is that it grades honestly. That claim is not compatible with a grading sheet that contains a field for arguing with the grade.

What we kept, and have kept since, is a separate review log where missed ratings are examined for systematic patterns — not to forgive individual misses but to find structural problems. If a certain stat type is being missed directionally in one sport, that shows up in the log and we look at the model. This is different from documenting that a specific player had a specific unusual night. The former is about the method. The latter is about the feeling.

Marcus made an observation during the cleanup that I wrote down because it seemed right in a way I hadn't articulated:

"The column was for us. The grade is for the record. We were mixing those up."
The distinction is simple and we had ignored it for four months.

I still don't know where the line is between documentation that improves a process and documentation that insulates a process from honest review. My working assumption is that any field which can be used to explain a miss without changing the grade is probably on the wrong side of that line — but I hold that loosely, because there are surely cases where the explanation matters and the grade is still correct. The question I haven't answered is how to tell them apart before you've spent four months finding out.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top