What We Do With a Rating We Can No Longer Grade
The grading window closes on a fixed schedule. Every rating we publish gets a result attached to it within a set number of days, or it gets flagged, or — and this is the case we spent most of last winter arguing about — it gets neither, because something happened between publication and resolution that made grading impossible. The player was ruled out after the rating was locked. The event was suspended and never completed. The stat category was quietly redefined by the source we used to build the corpus. In each of those cases, we had a rating sitting in the log with no grade attached, and no clean way to attach one.
That is not a rare edge case. Over a recent twelve-month stretch we counted forty-one ratings in that condition. Forty-one is not a large number in absolute terms, but it is large enough to matter for calibration, which is the only output we actually care about. If you are trying to measure whether your stated confidence matches your observed accuracy, a gap of forty-one unresolved records is not a rounding error. It is a structural problem that will quietly corrupt every calibration figure you publish, and it will do so in whichever direction flatters you most — which is the direction you should trust least.
We had been vague about this for longer than we should have been. The doctrine on our end has always been that you do not delete a rating, you grade it — but we had not been honest about what we did when grading was genuinely unavailable. The answer, for most of that period, was that we left the record open and moved on. That is not a policy. That is avoidance dressed up as patience.
Honest stories about the attempts, mistakes, deals, and numbers behind everyday hustles.
The Forty-One Records That Couldn't Be Closed
The problem splits into three distinct types, and we treated them as one type for too long.
The first is a pre-event disqualification: the rating was published, then the player was scratched or withdrawn before the event began. This one feels like it should be simple — no event, no grade — but it is not simple at all. If we published a rating at a high confidence level and the player was scratched four hours later, we still made a claim. The claim was that this player was expected to perform in a certain way, and we were confident enough in that expectation to publish it. The scratch does not retroactively make the claim correct. It makes the claim unresolvable, which is a different thing entirely.
The second type is a mid-event suspension: the game or match started, accumulated some statistics, and was then stopped and not resumed. We had three of these in the hockey corpus alone during one stretch. The player in question had logged some portion of their expected output before the stoppage. Do we grade against the partial stat line? Do we treat it as void? We tried both approaches at different times, and neither produced a calibration figure we felt comfortable publishing.
The third type is the one that caused the most internal friction: a retroactive stat revision. The source we used to build the corpus adjusted its historical figures — correcting an error, reclassifying a play type, updating a formula — after the rating had already been published and the grading window had already opened. The number the rating was built against no longer existed in the source. We have written about the specific damage that stat version drift can do to a corpus, but that piece was about building. This was about grading something that had already been built and published against a figure that had since been revised away.
How We Tried to Resolve the Unresolvable
Our first attempt was a category we called void with cause. The idea was to remove unresolvable ratings from the calibration denominator entirely, log the reason, and move on. Remy, who handles most of our grading infrastructure, built the tagging system in about a day. It worked cleanly from a technical standpoint.
The problem was what it did to the calibration figures. Once we started voiding records, our stated-confidence-to-observed-accuracy gap narrowed noticeably. The ratings we were most likely to void were the ones where something unusual had happened — a late scratch, a suspension, a stat revision — and unusual events cluster around the same conditions that make ratings hard to get right in the first place. We were, in effect, removing from the denominator a disproportionate share of the situations where we were most likely to be wrong.
"We built a process that made us look better at grading by removing the grades we couldn't do," Remy said, after running the before-and-after numbers. "That's not a process. That's a filter."
The second attempt was to assign a neutral grade — a fixed value representing neither a hit nor a miss — to every unresolvable record. The neutral grade would be included in the denominator but would contribute zero to the confidence-accuracy gap in either direction. This felt more honest than voiding. It also felt slightly arbitrary, because it was. We were manufacturing a grade for a record that did not have one, and calling it neutral did not make it less manufactured.
We also tried, briefly, grading mid-event suspensions against the partial stat line using a prorated threshold. If the event reached sixty percent completion and the player had accumulated sixty percent of the expected output, we counted it as a resolution. This was the attempt I am most embarrassed about. It introduced a second layer of estimation on top of the first one, and it did so in a way that was almost impossible to audit later. The principle we had written about before publication applies just as cleanly after it: if you cannot grade it cleanly, you have not actually graded it.
What the Ambiguity Actually Cost Us
The honest answer is that we do not know the full cost, and that is itself a cost.
What we can say is that our calibration figures for the two quarters in which we were using void with cause most aggressively are not trustworthy. We published confidence-accuracy summaries during that period that we now treat as suspect. We did not retract them — there was nothing factually false in them — but we added a note to the internal record that the denominator was trimmed in a way that likely favored our apparent accuracy. That note exists. Nobody outside the shop has seen it, because it lives in our grading log and not in the published pieces. That asymmetry bothers me more than the error itself.
The prorated threshold experiment cost us something different: it cost us a clean audit trail. When Remy went back through those records six months later to check our work, she could not always reconstruct which threshold we had used or why. We had been inconsistent in our application, which meant the grades were inconsistent, which meant the calibration figures built on them were inconsistent. The records looked resolved. They were not resolved. They were papered over.
There is also a subtler cost that I want to name carefully. The unresolvable records were not evenly distributed across stat types or sports. They clustered in specific conditions — high-variance events, late-breaking availability changes, sports where the corpus is thinner. Those are exactly the conditions where our ratings carry the most uncertainty to begin with. By handling those records inconsistently, we were being least rigorous precisely where rigor was most needed. The principle about distrusting results that please you applies here: the calibration figures that came out of that period were tidier than they had any right to be, and we should have been suspicious of that tidiness much earlier than we were.
The Policy We Actually Kept
What we settled on is not elegant. It is a set of three explicit rules that replace the earlier ambiguity, and each one costs us something.
Pre-event disqualifications are logged as unresolvable and included in the denominator with a fixed penalty applied to the confidence figure at which the rating was published. The penalty is not a grade. It is a discount applied to the stated confidence on the grounds that we published at a moment when the situation was less stable than our confidence implied. We are not claiming the rating was wrong. We are claiming that our confidence was not warranted by the available information, and we are recording that claim permanently.
Mid-event suspensions are voided only if the event reached less than a defined completion threshold before stopping. Above that threshold, we grade against the actual accumulated stat line with no proration. This produces some grades that feel unfair in individual cases — a player pulled from a match at seventy percent completion gets graded on seventy percent of the expected output, full stop — but it is consistent and auditable, which matters more than it feeling fair.
Retroactive stat revisions are the hardest case, and our rule is the bluntest: we re-grade the record against the revised figure, note the revision in the log, and accept that the grade may differ from what we would have produced at the time. The rating was built on a number that no longer exists in the source. Pretending otherwise does not make the record cleaner; it just makes the discrepancy invisible.
None of these rules make the unresolvable records go away. What they do is make our handling of them consistent enough that the calibration figures built on top of them mean something. Remy ran the calibration numbers under the new policy for the first full quarter of its operation and the confidence-accuracy gap was wider than it had been under the old approach. That is almost certainly the correct direction. A shop that publishes its own error rate should expect its error rate to look worse when it stops hiding part of it.
We still have not fully resolved what to do when a stat category is discontinued entirely — not revised, not corrected, but simply removed from the source with no replacement. We have two records in that condition right now, sitting open in the log. The policy above does not cover them. I am not sure any policy covers them cleanly, and I am genuinely uncertain whether the right answer is to manufacture a resolution or to leave them open indefinitely as a record of the limits of the method.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.