The Rating We Should Never Have Published
There is a specific kind of wrong that is worse than ordinary wrong. Ordinary wrong is a rating that failed because the information was incomplete, or the sample was thin, or the player had a bad night. That kind of wrong is just noise, and the method accounts for it. The kind I am describing here is a rating that cleared every gate cleanly, earned a high confidence label, was graded against what actually happened, and was wrong in a way that the corpus had already told us about — if we had been willing to read it that way. We were not willing. That is the whole story, and I am going to tell it in full because that is what this desk is for.
It happened during a stretch of basketball evaluation that had been going unusually well. Calibration was tighter than normal. The grading window had closed on four consecutive weeks without a significant miss. I think that is relevant context, not because a good run should change the method, but because it did change the temperature in the room. When things are going well, the gates feel like formalities. They are not formalities. They are the method.
The player in question — I will call him the candidate, because we do not attach names to directions here and because the mechanism is the lesson, not the individual — had a corpus that looked, on the surface, like one of the strongest we had assembled for a basketball evaluation that season. Volume of games, consistency of output, no red flags at any of the standard disqualification points. The rating came out high. We published it. Then we graded it.
Practical guides for test analytics, reliability, observability, reporting, and AI-driven quality.
The Corpus Looked Clean Because We Stopped Reading It Early
The candidate had played a large number of games across two seasons, and his per-game output in the primary statistic we were rating was remarkably stable. Stable is good. Stable is what we want to see. But stability in aggregate can hide a split that matters enormously, and we did not look for the split until after the rating had already been published and graded.
The split, when we finally pulled it, was between games where the candidate was operating as the primary option and games where he was not. In roughly a third of his appearances — appearances we had included in the corpus without flagging — he had been deployed in a secondary role due to lineup configurations that were no longer in place. His numbers in those games were substantially lower. We had averaged them in anyway, which pulled his baseline up toward the primary-role games without our ever acknowledging that the role distinction existed.
This is the kind of thing that what counts as enough history actually means in practice. It is not only about volume. It is about whether the volume you have is measuring the same thing throughout. Ours was not. We had a large sample of two different players and had treated it as a large sample of one.
"The corpus looked deep because it was long. Long and deep are not the same thing. We should know that." — Renata
She said it the day after the grade came back, and she was right, and it was not a new observation. We had written about this exact failure mode before. Knowing about a failure mode and catching it in the moment are different skills, and we are better at the former than the latter.
How the Rating Cleared Every Gate Anyway
I went back through the gate log after grading. The candidate had cleared availability, cleared the activity recency check, cleared the expected-role confirmation, and cleared the sample-size minimum. Every gate passed. The problem is that the expected-role gate was answered with reference to the current lineup configuration, which was indeed the primary-role setup — so the gate passed correctly, on the information available at the time. What the gate did not do was ask whether the corpus reflected that role, or a mixture of roles, or something else entirely.
That gap is not new. We have written about how most candidates should die before scoring, and the argument there is that the gates exist precisely to catch situations where the surface information looks clean but the underlying data is compromised. The expected-role gate passed because the role was confirmed. It should also have asked: is the corpus role-consistent? It did not ask that. It does now.
The line reading also gave us nothing to work with — or rather, it gave us something we interpreted as confirmation when it was not. The line was set in a range consistent with the candidate's aggregate numbers, which made sense given that whoever set it was probably looking at the same aggregate we were. When everyone agrees, that convergence can feel like validation when it is actually just two parties making the same unexamined assumption. We noted the agreement. We did not ask what assumption it was built on.
What the Grade Said and What We Had Said
The rating carried our second-highest confidence label. That label means, in calibration terms, that we expect to be right at a rate commensurate with that confidence — not always, but across enough instances that the label means something. When we closed the grading window, this one was wrong. That is one data point, and one data point does not break a calibration record. But the reason it was wrong is what concerns me.
A rating that fails because of bad luck, or an unpredictable event, or genuine noise in the system — that is not a calibration problem. It is expected attrition. A rating that fails because the corpus was measuring the wrong thing, and we knew enough about corpus construction to have caught it, and we did not catch it — that is a calibration problem of a different kind. It means the confidence label was not earned. We said we were confident because the surface looked clean, not because we had done the work that confidence requires.
The cost, then, was not just one wrong rating. It was a small corruption of the confidence label itself. If the label can be assigned when the corpus has not been interrogated properly, then the label is measuring our comfort level, not our accuracy. Those are not the same thing, and we built this shop on the premise that we would not confuse them.
We also did not catch it during the shrinkage step, which should have pulled the estimate back toward the base rate more aggressively given the sample heterogeneity. It did not, because the heterogeneity was not visible at that stage — it had been averaged away earlier. Shrinkage cannot correct for a problem it cannot see. That is a sequencing issue, and it is on us.
What Changed in the Gate Sequence After This
We added a role-consistency check to the corpus review step. It is not a gate in the strict sense — it does not automatically disqualify a candidate — but it flags any corpus where more than fifteen percent of appearances involve a materially different deployment than the current expected role. When the flag fires, the reviewer has to make an explicit decision about whether to trim the corpus, reweight it, or proceed with a notation. The decision has to be recorded. Before this, there was no mechanism forcing that question to be asked.
We also changed the language on the confidence labels internally. The second-highest label now carries a parenthetical in our grading log: corpus role-consistent confirmed. If the confirmation is not there, the label cannot be assigned. It is a small bureaucratic friction, and Tomás complained that it would slow things down during high-volume periods. He was right that it would slow things down. I did not find that compelling.
The rating itself was graded and kept in the record, as everything is. I want to be precise about what that means: the rating is not in the record as a cautionary footnote. It is in the record as a data point in our calibration history, counted the same way every other rating is counted. The fact that it was wrong for an identifiable reason does not give it special status. It just makes the postmortem easier to write.
What we did not change is the underlying confidence in the corpus-first approach. The failure here was not that we built a corpus — it was that we built one carelessly and then trusted it completely. Sample size still beats recency, and a large body of prior performance still outranks a short hot streak. But sample size only beats recency if the sample is measuring what you think it is measuring. That qualifier was always implicit. After this, we say it out loud.
The thing I keep returning to is that the candidate's numbers in his primary role were, in fact, quite good — probably good enough to have earned a rating on their own, trimmed corpus and all. We might have arrived at a similar conclusion through more honest means. Or we might not have. I genuinely do not know, and I find that uncertainty more useful than the false certainty we published.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.