A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Window We Agreed On and Ignored

We wrote the rule down. That is the part that still bothers me. It was not an informal understanding or something we meant to formalize later — it was a typed sentence in the method document: every rating is graded at the close of the agreed window, not before, not after, not when convenient. We had written it specifically because we had already been burned once by setting windows after we saw the results, and we were not going to do that again. The document existed. The sentence was in it. We ignored both.

What happened is not complicated. A cluster of basketball ratings from a particular stretch of the season were sitting in the grading queue, and one of them was looking poor about two-thirds of the way through its window. Not catastrophically poor — the player had simply been quieter than the corpus suggested he would be, and the underlying number was drifting toward a miss. Someone noticed. Someone mentioned it in the log. And then, instead of waiting, we started discussing whether the window itself had been set correctly.

That discussion is the tell. When the rating was tracking well, nobody questioned the window. The moment it looked like a miss, the window became a methodological concern. That asymmetry is exactly the thing we are supposed to be guarding against, and we walked straight into it anyway.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

Why the window started moving the moment a rating went soft

The grading window we use for player ratings is not arbitrary. It is set before evaluation begins, based on the type of event, the stat category in question, and how many opportunities the player is expected to have within the period. A basketball assist rate over a four-game stretch means something different from the same rate over twelve games, and the window is supposed to reflect that. We document the window at the time we issue the rating. The documentation is timestamped.

What the documentation cannot do is stop us from reading it later with motivated eyes. The basketball rating in question had a window of eleven games. At game seven, the player had produced in a way that put him below the threshold on which the rating was based. Nora pulled the original corpus entry and pointed out, correctly, that seven games was within the variance we had already measured — that the corpus showed this player going quiet for stretches of exactly this length before recovering. The rating was not obviously wrong. It was just uncomfortable to look at.

The problem was that Nora's observation, while accurate, arrived at the wrong moment. We were not reviewing the corpus because we had scheduled a mid-window check. We were reviewing it because the rating was soft and we wanted a reason to feel better about it. The distinction matters enormously for calibration. A check that happens because a rating looks bad is not a neutral methodological review — it is a search for exoneration, and it tends to find what it is looking for.

"We didn't change the window in the document," Nora said afterward. "We just stopped treating the document as the authority. Which is worse, actually, because at least a changed document leaves a record."

She was right. What we did was more insidious than a clean edit. We let the window sit unchanged in writing while behaving as though it no longer applied. The rating got an informal extension — not recorded, not justified prospectively, just quietly granted because closing the window on schedule would have meant logging a miss.

The informal extension we gave ourselves and called a review

The language we used to justify this was methodological. We told ourselves we were examining whether the original window had been appropriate for this stat category — whether eleven games was really the right period for an assist-rate rating on a player whose role shifted depending on lineup context. That is a legitimate question. It is also a question we should have asked before the window opened, not at game seven when the rating was soft.

We did eventually close the window. We waited until game eleven as originally documented, and the player did recover somewhat — not enough to flip the rating to a clear hit, but enough that the outcome landed in the ambiguous zone where we could argue either way. We argued the better way. We logged it as a marginal pass.

I want to be precise about what was wrong there. The outcome at game eleven was real. We did not fabricate the number. But the decision to wait until game eleven — rather than close at game seven when we first started questioning the window — was not driven by the method. It was driven by the fact that game eleven produced a more favorable result. We did not know that at game seven, but we behaved as though we suspected it, and that suspicion shaped the process. Window movement that protects a rating is exactly the failure mode we had catalogued before. We catalogued it and then reproduced it.

The attempt to frame this as a legitimate review also contaminated the broader grading session. Two other ratings from the same stretch were graded on schedule without incident. In retrospect, the consistency of those two gradings makes the basketball rating look worse by contrast — if the window question was genuinely methodological, it should have applied to all three.

What the informal extension actually cost the calibration record

The immediate cost was a logged pass that should probably have been a logged miss. That matters less than it sounds — one rating in a grading session is not enough to move the calibration numbers meaningfully. The larger cost was what it did to the window rule itself.

Once a rule has been broken once without consequence, it is no longer quite a rule. It is a guideline with an exception, and exceptions have a way of multiplying. In the two months after this episode, we reviewed the grading log and found three other instances where windows had been implicitly extended — none as deliberate as the basketball case, but all of them showing the same pattern: the extension happened when the rating was soft, never when it was tracking well. We had not noticed because none of the extensions were documented. They existed only in the gap between what the method document said and what we actually did.

This is the kind of error that does not show up in a hit rate. If you only count outcomes, a marginal pass looks identical to a clean pass. The calibration problem only becomes visible when you ask whether your stated confidence at the time of issuing the rating matches your observed accuracy — and that calculation requires honest grading, which requires honest windows. A window that is not fixed in advance is not a window at all; it is a retrospective judgment dressed in procedural clothing.

We also lost something harder to quantify: the ability to trust our own grading log as a clean record. Every calibration check we run now has to carry a small asterisk — these numbers assume the grading was honest, and we know at least some of it was not. That uncertainty is not fatal, but it is real, and we put it there ourselves.

What we actually changed in the method after this, and what we did not

We added a gate. Before any mid-window corpus review can happen, someone has to document the reason for the review in the log, and the reason cannot reference the current trajectory of the rating. "The player's role has shifted due to a lineup change" is a valid trigger. "The rating is tracking soft" is not. The log entry is timestamped and cannot be edited after the fact — we use a append-only format precisely because we do not trust ourselves to leave edits alone.

We also went back and regraded the basketball rating under the original schedule. We closed it at game seven with the numbers as they stood then. It logged as a miss. We updated the calibration record accordingly. This meant our stated confidence on that rating type was now running slightly above our observed accuracy — a small calibration gap, but a real one, and one we would not have found if we had left the favorable grade in place.

What we did not change is the window length itself. Eleven games for an assist-rate rating on a role-dependent basketball player is probably right, and revising it downward because one rating went badly would have been another version of the same error. Nora argued for a review of all window lengths across stat categories, which is a reasonable thing to do periodically — but doing it immediately after a miss would have looked, and probably been, motivated. We scheduled it for the following quarter instead, when the specific rating would be far enough in the past to stop mattering.

The grading column itself also got an audit. We had been resetting windows after embarrassing misses more often than I had realized — not always consciously, but the pattern was there in the timestamps. Cleaning that up took most of an afternoon and produced a calibration record that was noticeably less flattering than the one we had been looking at. That is probably the most useful thing that came out of this episode: not the new gate, but the honest count.

The rule was written down. We still broke it. I am not sure what follows from that — whether better documentation would have helped, or whether the problem is something that documentation cannot reach, something about how uncomfortable it is to watch a rating fail in real time and do nothing. The gate we added addresses the mechanism. Whether it addresses whatever was underneath the mechanism is a different question, and one I am less confident about.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top