A Sport We Gave Up On, and Why
We dropped a sport. Two seasons of work, a derived parameter set, a working pipeline, and we turned it off and have not gone back.
The reason had nothing to do with accuracy. Our accuracy there was unremarkable in both directions. The problem was that we could not reliably tell what had happened afterwards.
The origins and reasoning behind familiar things.
Outcomes That Arrived Late, Amended, or Not at All
Grading requires an authoritative account of what occurred. In our other sports that account is available within hours and is stable once published.
In this one it was published quickly and then amended — sometimes days later, occasionally more than once. Individual contributions were reclassified after the fact for legitimate procedural reasons. And a meaningful minority of events never produced a clean record at that level of detail at all.
So we had three categories of rating: graded correctly, graded and then invalidated by an amendment, and ungradeable. The third category was large enough to matter and — this is the part that killed it — it was not random. Events with incomplete records skewed toward particular circumstances, which meant the subset we could grade was a biased sample of the ratings we had made.
A biased grading sample is worse than no grading. It produces a confident number about a systematically unrepresentative slice, and every calibration conclusion drawn from it is quietly wrong in a direction you cannot measure.
It is worth being precise about why a biased grading sample is uniquely bad rather than merely inconvenient. An unbiased gap costs you precision — your numbers are noisier and you need more of them. A biased gap costs you validity, which no amount of additional volume repairs. We could have run that sport for a decade and accumulated an enormous, confident, systematically misleading record of our own performance, and nothing inside the system would ever have flagged it. Volume actively hurts you here: it makes the wrong number look sturdier.
Three Attempts to Grade Around the Problem
We tried extending the window, so amendments would land inside it. That helped with amendments and did nothing for the missing records, and a longer window has its own cost — ratings sit unresolved, and the feedback that makes the method improve arrives later.
We tried treating ungradeable ratings as a category and reporting them separately, honestly, as a known gap. That was the most defensible option and we ran it for a season. It did not work either, for a reason I did not anticipate: nobody, including us, could hold the caveat in mind while reading the numbers. The graded figure was right there and the disclaimer was a sentence underneath it.
We tried restricting coverage to the subset of events with reliable records. That worked technically and eliminated most of the value, because the reliable subset was the well-covered, heavily-watched end — exactly where an evaluation shop has least to add.
“We can rate it, we just can't mark our own homework,” Priya said. “I don't know what we're doing if we can't do that.”
What Turning It Off Cost, Including the Part That Was Pride
Two seasons of derivation work, a parameter set nobody will use, and roughly a fifth of our coverage by volume.
It also cost a genuinely held belief. I had assumed our constraint was analytical — that the hard part of this work was evaluating well. It is not. The hard part is being able to check yourself, and a domain where checking is unreliable is a domain where we have no way to earn the confidence we would be expressing.
I resisted the decision for about four months, and my argument was that our accuracy was fine. That argument was irrelevant and I made it repeatedly. Accuracy measured on a biased sample is not accuracy, and I knew that in every other context.
Gradeability Became an Entry Requirement
We now assess whether a domain can be graded before we assess whether it can be modelled. Specifically: how quickly does an authoritative record appear, how often is it amended, and what proportion of events produce no usable record — and is that proportion random.
That last question is the one that decides it. A gap you can characterise is survivable. A gap that correlates with the circumstances of the event is not, because it makes your own scorecard unreliable in a way no caveat repairs.
The doctrine entry is four words: if you cannot grade it, do not publish it. It has stopped us adding two domains since, both of which looked interesting and neither of which we regret skipping.
The parameter set is still in the repository, disabled, with a note. Somebody proposes turning it back on about once a year, and the note is usually enough.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.