A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Window We Kept Moving

The rating in question was for a basketball player we had tracked across two full seasons. The corpus was deep — deeper than most of what we work with — and the signal had survived every gate we put in front of it. We were, by our own internal accounting, confident. Not loudly confident. The quiet kind, which is worse, because it takes longer to notice.

The problem surfaced during a routine calibration pass. The grading window we had specified in advance was thirty days from the date of publication. On day thirty-one, the player had not yet played enough games to return a clean result — an injury absence, then a managed return, then a schedule quirk that compressed everything. We made a note. We extended the window to forty-five days. The result still looked inconclusive. We extended it again.

By the time we stopped, the window had moved four times. The rating had never changed. Only the deadline for judging it had.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

Why a Moving Window Is Not a Grading Window

There is a version of this story where we look reasonable. Injuries happen. Schedules compress. A thirty-day window that catches only eight games instead of the expected eighteen is genuinely less informative than one that catches twenty-two. We told ourselves this for a while.

The version that is actually true is less flattering. We had a number we liked, and we kept adjusting the measurement instrument until the number could survive. That is not grading. That is curating, which is the thing we exist to not do. The grading window has to be fixed in advance precisely because the moment you know the outcome direction, you lose the ability to choose a window neutrally. We knew the outcome direction. We were not neutral.

Marta, who handles most of our calibration logging, put it plainly when she reviewed the audit trail:

"The four extensions are all documented, and every one of them has a reasonable-sounding note next to it. That's actually the tell. A window you move once for a genuine data gap looks like one thing. A window you move four times, each time with a tidy explanation, looks like something else."

She was right. The explanations were not fabricated — the injury was real, the schedule compression was real — but the pattern of reaching for an explanation each time the window came due was something we had manufactured.

What We Tried When We Finally Noticed

Once Marta flagged the pattern, we ran the rating against all four windows simultaneously and recorded a result for each. The idea was that if the extensions were genuinely neutral — if we had been moving the window for data-quality reasons rather than outcome-protection reasons — the results across the four windows should cluster. Inconclusive data is inconclusive regardless of when you measure it.

They did not cluster. The thirty-day result was a miss. The forty-five-day result was marginal. The sixty-day result was a narrow pass. The final extended window, at seventy-eight days, was the only one that returned something we could have called a clean confirmation. We had, without quite deciding to, selected the one window that made us look correct.

We also went back through the preceding six months of ratings to check whether this was an isolated case or a habit. We found two other instances where the window had moved once — both documented, both with plausible notes — and one where it had moved twice. None were as egregious as the basketball case, but the direction was consistent: windows moved when ratings were under pressure, and stayed fixed when they were not. That asymmetry is the signature of a process protecting itself.

This connects to something we had already written about in a different context: we never delete a rating, only grade it. The rule exists because deletion is the most obvious form of outcome protection. What we had found was a subtler version of the same instinct — not erasing the record, but softening the court it had to play on.

What the Moving Window Actually Cost Us

The immediate cost was a calibration record that overstated our accuracy by a small but real margin. A rating that should have been logged as a miss was logged as a pass, and that pass fed into the confidence estimates we use when we decide how hard to shrink future ratings in the same category. In other words, the inflated record made us slightly less aggressive about pulling subsequent estimates toward the base rate. The error propagated quietly.

The larger cost was harder to quantify. Calibration — the match between stated confidence and observed accuracy — is the only output we actually care about at this shop. Being right 70% of the time when you said 70% is the whole game. A single manipulated window does not destroy that record, but it introduces a kind of rot: if the grading process can be adjusted after the fact, then the confidence figures we publish are not confidence figures. They are something more like wishes with decimal points attached.

We had also, without meaning to, built a small internal precedent. The people who logged the extensions had each done so individually, each with a reasonable note, and none of them had flagged it as a pattern because none of them had seen the full sequence. The process had failed not because anyone was dishonest but because the audit trail was not designed to surface this specific shape of error. We were updating something without fully knowing why, and the log was not asking the right questions.

What the Process Looks Like Now, and What We Are Still Not Sure About

We made two structural changes. The first is simple: window extensions now require a second signature — someone who was not involved in the original rating has to agree that the extension is for a data-quality reason and not an outcome-adjacent one. In practice this means Marta reviews every extension request before it goes into the log. She is not always available and the process has already slipped once, which tells us the rule is right but the enforcement is still soft.

The second change is that we now run the original window result in parallel with any extended result, and both appear in the calibration record. If a rating passes on the extended window but fails on the original, it is logged as a conditional pass with a flag — it counts neither as a clean hit nor as a miss, but it does not disappear. This is uncomfortable because it produces a category of result that is hard to aggregate, and our calibration calculations were not designed for ambiguity. We have been living with that discomfort for about four months now and have not resolved it cleanly.

What we did not change: the thirty-day default. We considered shortening it to twenty-one days on the theory that a tighter window would leave less room for manipulation. The counterargument, which won, was that a shorter window would simply produce noisier results and give us more legitimate reasons to extend. You cannot fix a judgment problem with a shorter ruler.

The basketball rating, for the record, was re-logged under the original thirty-day window as a miss. It is in the record now, unqualified. The player's subsequent performance, over the following season, was more or less what the rating had predicted — which is the kind of thing that would have felt vindicating if we had not already spent several months demonstrating that we could not be trusted to grade our own work on that one.

The question we have not answered is whether any window-extension policy can be made manipulation-resistant, or whether the problem is simply that the people doing the grading are the same people who published the rating. A shop that grades its own output is always going to have this tension. We have not found a structural fix for that, and we are not sure one exists.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top