PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

If You Cannot Grade It, Do Not Publish It

For about four months in our second year, we published ratings we could not grade. Not because we were hiding them — we told ourselves we would grade them later, once the situation resolved, once the data caught up, once we had a cleaner way to define the outcome. The situations did not resolve. The data did not catch up. What we had, at the end of those four months, was a section of the record that was technically published and practically invisible to the calibration process. We had laundered roughly sixty ratings out of the error count without meaning to.

The mechanism was mundane. Certain categories — a player returning from a layoff with no clear minutes baseline, a hockey skater whose role had shifted mid-season in ways the box score did not capture — resisted clean grading. The outcome happened. The number existed. But connecting the outcome to the rating required judgment calls that changed depending on who made them and when, and we were a small shop with inconsistent bandwidth. So the ratings sat in a queue, aged past usefulness, and were quietly not counted. We did not delete them. We just never closed them. The distinction felt meaningful at the time. Looking back, it was not.

The rule that came out of that period is the bluntest one we have: if you cannot specify, before publication, exactly how a rating will be graded, you do not publish the rating. Not "we will figure it out." Not "probably by outcome X." Exactly how, in writing, before the event begins. The rule has cost us material we thought was good. We have kept it anyway, though I want to be precise about why — because the reason is not moral, and it is not obvious.

Wanderlust y Couture with LuxeSofia

Discover luxury hotels, chic city stays, and beautiful escapes around the world.

Learn more

The Ratings That Lived Outside the Count

The problem was not the ratings themselves. Several of the sixty were probably fine — reasonable estimates built on adequate corpora, cleared through the gates, shrunk toward the base rate the way they were supposed to be. The problem was that they existed in a state where they could not embarrass us, and a shop whose entire premise is that it publishes its own error rate cannot have a category of output that is structurally protected from embarrassment.

Riya noticed it first. She was running the quarterly calibration pass and flagged that our stated confidence on the ungraded queue was, on average, higher than on the graded pool. That is almost certainly not a coincidence. When grading is uncertain, we were apparently more willing to publish a confident estimate — possibly because the uncertainty felt like it lived in the grading mechanism rather than in the rating itself. It did not. Uncertainty about how to measure an outcome is uncertainty about the outcome, and we had been treating it as a separate problem.

This is the kind of thing that accuracy and calibration measure differently. A high hit rate on the graded pool told us nothing about whether the ungraded pool would have dragged that number down. We did not know. That was the point. We had created a shadow record that was neither honest nor dishonest — it was just absent, and absence in a calibration log is its own kind of lie.

The Grading-Criteria Draft We Wrote After the Fact

The first fix was to go back and grade the sixty retroactively. We spent two weeks on it. We wrote explicit criteria for each category that had resisted grading, applied them consistently across the queue, and closed every open rating. The hit rate on the retroactive batch was 51 percent on estimates where we had stated roughly 68 percent confidence. That is a bad number. It is also almost certainly not the real number, because the criteria we wrote after the fact were written by people who had seen how things turned out.

I do not think we cheated deliberately. I think it is nearly impossible to write grading criteria in full knowledge of the outcome and have those criteria be genuinely neutral. The human brain does not work that way, and ours did not either. The retroactive pass gave us a closed record, which was better than an open one, but it did not give us a trustworthy record. We noted this in the calibration log and moved on, which is the correct thing to do and also unsatisfying.

The second fix was a pre-publication checklist: before any rating went out, the analyst responsible had to write, in a single sentence, the exact grading condition. "This rating grades as correct if the recorded outcome at full time is within N of the estimate." Something that specific, something that could be checked by someone who had not seen the rating before and did not know what we expected. If the sentence could not be written, the rating did not publish. We ran this for six weeks before we found the edge cases.

What the Checklist Did Not Catch, and One Week It Made Worse

The checklist worked on the obvious cases. It failed on the structural ones. There is a category of player — common in individual sports, present in team sports too — where the grading condition depends on a contextual variable that is not known until after the event begins. A tennis player whose match format changes. A baseball pitcher whose role is altered in the late innings in a way that changes what the counting stat means. We had written grading conditions for these that were technically specific and practically useless, because they assumed a stable context that the event did not provide.

The week this became expensive is documented elsewhere — the week everything was confidently wrong involved, among other things, three ratings that graded correctly by our written criteria and incorrectly by any honest reading of what we had predicted. We had written the criteria tightly enough to pass the checklist and loosely enough to survive outcomes we had not anticipated. That is a form of the same problem we started with, arrived at from the opposite direction.

"The checklist told us we had a grading condition. It did not tell us the grading condition was any good. Those are different gates." — Riya

She was right, and we adjusted the process again: the grading condition now has to survive a five-minute adversarial review by a second person, whose job is specifically to find scenarios where the condition would produce a misleading result. This slows publication. It has killed ratings we thought were solid. It has also, twice, caught conditions that would have closed as correct on a technicality and incorrect in fact. We kept the step.

What the whole episode cost us, beyond the two weeks of retroactive work and the ratings we could not publish, was a cleaner story about our early record. We do not delete ratings, we grade them — that is the doctrine — but grading them badly and grading them well are not the same act, and our first two years contain a mixture of both. The calibration numbers from that period are published with a footnote. The footnote is honest. It is also a reminder that a rule written after a problem is always at least one problem behind.

The Version of the Rule That Has Lasted

The rule as it stands now has three parts. First: the grading condition must be written before the event begins, in language specific enough that someone with no prior knowledge of the rating could apply it. Second: the grading condition must survive adversarial review. Third: if the event produces a context the condition did not anticipate — a format change, a role shift, an abandonment — the rating is graded as void, logged as void, and counted in the void rate, which is itself a number we track and publish. A void is not a miss and it is not a hit. It is a third outcome, and pretending it does not exist is what got us into the original problem.

The void rate has been informative in ways we did not expect. It is highest in individual sports, which turns out to be a consistent pattern — individual sports break our method in several directions, and contextual instability is one of them. It is also higher early in a season than late, higher for players returning from absences, and higher in sports where substitution rules give coaches meaningful discretion over how a player's counting stats accumulate. None of that was obvious before we started tracking it. It became obvious quickly once we did.

The deeper thing the rule changed was what we were willing to evaluate at all. Before the checklist, a candidate could pass the gates and reach publication even if the outcome was genuinely ambiguous — we told ourselves the ambiguity was in the grading, not the rating. After the rule, the inability to write a clean grading condition became a gate in itself. Most of the candidates who died at that gate were ones that should have died earlier anyway. The grading condition was just a later-stage symptom of the same underlying problem: not enough stable signal to say anything precise.

We publish fewer ratings now than we did in the first two years. The calibration numbers are better, though I am cautious about claiming that as a success — a smaller, cleaner sample can look well-calibrated for reasons that have nothing to do with the method improving. That is a different problem, and we are probably not done with it.

The rule is simple to state and expensive to keep, which is usually the sign that it is doing something real. What I am less sure about is whether the adversarial review catches the right failure modes, or whether it mostly catches the ones we already know to look for — which would mean the conditions we are most confident in are still the most dangerous ones.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top