The Grading Window We Reset After the Miss
The miss itself was not spectacular. A basketball player we had rated highly on rebounding volume put up roughly half his expected total across three consecutive games, and the corpus said that was unusual enough to register. We graded it, it failed, and the failure went into the log. That part was routine. What was not routine was what we did next: we quietly extended the grading window by eleven days, let the player's numbers recover, and reclassified the rating as still active. Nobody outside the shop saw the original close. We did not announce the extension. We just moved the date and kept going.
I am writing this up because the shop's entire premise is that it publishes its error rate, and an error rate that excludes the errors we found embarrassing is not an error rate at all. It is a press release. The miss itself would have been a footnote. The window reset is the actual story, and it took us longer than it should have to say so plainly.
What made it worse was that we had already written, at some length, about why the grading window has to be fixed in advance. We knew the rule. We had articulated the rule. We broke the rule anyway, and we broke it in the specific way that the rule exists to prevent: we saw the result first, then adjusted the frame around it.
See the cash truly available after bills, payroll, taxes, and reserves before making your next move.
The rating that made us flinch before we graded it
The player — invented, as all our examples are — had a strong corpus behind him. Roughly two and a half seasons of per-game rebounding data, consistent minutes, no significant injury history in the window we were using. The gates had cleared him without argument. When the line came in lower than his trailing average, we read it the way we usually do: as a claim about what the market believed, not as an instruction. The rating came out confident. Confident enough that when the three-game stretch went badly, the failure was visible to anyone looking at the grading log.
That visibility is what moved us. Not the miss itself — we miss, and we record it. What moved us was the size of the gap between stated confidence and observed outcome. We had expressed something close to strong conviction on a rebounding estimate, and the player had not come within range of it across three separate games. The calibration implication was uncomfortable: if you say you are highly confident and you are wrong three times running, the honest reading is that you were overconfident, and overconfidence is the specific failure mode the shop was built to catch.
So instead of recording the overconfidence and moving on, someone — and I will not pretend it was not me — suggested that three games might not be a fair sample for a player with a two-and-a-half-season corpus behind him. That is not a wrong observation. It is, in fact, sometimes correct. The problem is that we made it after we had already seen the three games fail.
How we talked ourselves into extending the window
The argument we made internally was coherent on its surface. A player with a large, stable corpus should not be re-evaluated on the basis of three games. Sample size beats recency — that is the shop's oldest stated belief, and we leaned on it hard. We told ourselves we were being rigorous, that snapping a window shut after three bad games was the reactive thing to do, and that extending it was the patient, principled thing to do.
Remi, who handles most of the calibration review, put it more directly than I had been willing to:
"We were using 'sample size beats recency' as a reason not to close the grade. But sample size beats recency is a rule about building ratings, not about protecting them once they've already been tested. We flipped the doctrine around to face the other direction."
She was right. The doctrine about sample size is upstream of grading — it governs how much history we require before a rating is issued. It says nothing about how long we are allowed to keep a grading window open after the outcome has already occurred. Those are different questions, and we had deliberately blurred them because blurring them was convenient. The eleven-day extension let the player's numbers recover to something closer to his corpus average, and the rating ended up graded as a narrow miss rather than a significant failure. A narrow miss is logged and forgotten. A significant failure triggers a calibration review. We had, in effect, engineered our way out of a calibration review.
What the reset actually cost us in the calibration log
The immediate cost was small and the delayed cost was larger. In the short term, one rating moved from the "significant failure" column to the "narrow miss" column. The calibration numbers for that month looked slightly better than they should have. Nobody outside the shop noticed, because nobody outside the shop has access to our grading log at the daily resolution we keep internally.
The delayed cost was that we had introduced a decision rule we had not written down: that windows could be extended when the corpus was strong and the sample of failures was small. That unwritten rule then appeared, quietly, in two subsequent ratings over the following six weeks. Not extended by eleven days — extended by three or four days each time, which felt modest and therefore felt defensible. By the time we caught it in a process review, we had a de facto policy of post-hoc window adjustment that none of us had ever explicitly agreed to. We had each assumed someone else had sanctioned it.
This is the specific failure mode that moving a grading window to protect a rating produces: not one big dishonest act, but a series of small adjustments that each feel reasonable in isolation and are collectively corrosive. The calibration log is only useful if the grades inside it were produced by a consistent rule. Once the rule becomes "extend when it feels like the corpus supports it," the log is measuring our comfort with outcomes rather than our accuracy at predicting them.
We also lost something harder to quantify. The shop's credibility with itself depends on the grading log being honest. When the people maintaining it know that certain entries were produced by adjusted windows, the log starts to feel like a document we manage rather than a document we read. That is a different kind of damage, and it does not show up in any calibration percentage.
What we changed after we wrote it down honestly
The first thing we did was regrade the original rating using the original window. The significant failure went back into the log where it belonged. The calibration numbers for that month got worse. We noted the correction and the reason for it in the log itself, which is now standard practice: any retroactive change to a grading entry requires a written explanation attached to the entry, visible to anyone reviewing the corpus later.
The second thing we did was formalize what had been an implicit assumption. The grading window is now written down before the rating is issued — not as a general policy, but as a specific date attached to a specific rating at the moment of issuance. This is not a new idea; writing the threshold down before seeing the data is a principle we had already applied to other parts of the method. We had just failed to apply it consistently to the window itself. The fix was not clever. It was tedious and obvious, which is usually the sign that it will actually hold.
We did not add a rule prohibiting window extensions entirely. There are legitimate reasons to extend a window — a game postponed by weather, a player held out for a non-injury reason that was not public at issuance — and a blanket prohibition would create its own distortions. What we prohibited, specifically, is extending a window after the original close date when any portion of the graded event has already been observed. The distinction matters: a postponement before the event is a scheduling fact; an extension after three bad games is a judgment call made in the wrong direction.
Remi added a second check to the calibration review cycle, specifically looking for entries where the window close date does not match the date recorded at issuance. It fires about once a month and has caught two legitimate postponement adjustments and one case where someone had simply mistyped a year. Nothing suspicious since the policy change, which either means the policy is working or means we have gotten better at not noticing when we bend it.
The part I keep returning to is that the original miss was not that bad. We were overconfident on a rebounding estimate for a player with a strong corpus, and we were wrong. That happens. The interesting question is whether the instinct to protect a confident rating from a small sample of failures is a bias we introduced, or whether it is something reasonable that just needs a harder boundary around it — and I am genuinely not sure the answer is the same in every case.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.