A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Window We Set After Seeing Results

The grading window is supposed to be the easy part. You rate a player, you wait for the relevant period to close, you record whether the rating held. The discipline is not in the math — it is in the timing. You fix the window before the results exist, or you are not grading anything. You are just choosing which outcomes to count.

We did not always do this. For roughly the first eighteen months of formal grading, we set the window after we had already seen enough of the results to have a sense of how things had gone. We told ourselves we were being careful — waiting until the sample was "mature enough to be meaningful." What we were actually doing was letting the outcomes inform the frame. The distinction felt technical at the time. It was not technical. It was the whole thing.

This piece is about what that cost us, what we changed, and the specific moment when Renata sat down with the calibration logs and showed me, in about four minutes, that our hit-rate numbers were not measuring what we thought they were measuring.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

How a window set in hindsight flatters every number inside it

A grading window is a commitment: this rating covers this span of activity, and at the end of that span, we record a result. The commitment has to be made before the span ends. That sounds obvious written down. It was not obvious in practice, because the temptation to adjust the window is not dishonest in its motivation — it comes from a genuine concern about noise.

A basketball player misses two games with a minor injury inside your window. Do you extend the window to capture a cleaner sample? A hockey player is playing on a line that gets reshuffled halfway through. Do you close the window early, before the reshuffle contaminates the read? Both adjustments feel methodologically defensible. Both adjustments are, in effect, you editing the test after you have seen some of the answers.

The problem compounds because the adjustments are not random. We were not equally likely to extend a window when things were going well and when they were going badly. We extended windows when the early results looked noisy and unfair to a rating we felt confident in. We closed windows early when a clean run of confirming results had already appeared. The bias was not malicious. It was not even conscious for most of it. But it was systematic, and systematic bias in the grading window means every calibration figure downstream is wrong in the same direction.

We had written, in an earlier piece, about why the window has to be fixed before you look. We had written it as doctrine. We were not following it.

What we tried when Renata flagged the calibration gap

Renata had been running a side audit — not because anyone asked her to, but because she had noticed that our stated confidence levels and our observed accuracy were drifting apart in a way that did not feel like ordinary variance. She pulled eighteen months of grading logs and coded every window by whether it had been set before or after the first result in that window was recorded.

"About forty percent of the windows had at least one result already in the log before the window was formally closed. That's not a rounding error. That's a design problem."

She was right, and the number was worse than I expected. Forty percent is not a few slips. It is a practice. What we called "grading" was, in a meaningful share of cases, retrospective selection dressed up as prospective evaluation.

The first thing we tried was a soft fix: a rule that the window had to be logged in the system before the player's next relevant event, regardless of whether we had seen any results yet. This was better than nothing. It closed the most egregious cases — the windows we had been sitting on for weeks while results accumulated. But it did not fix the subtler problem, which was windows that were nominally set in advance but whose length had been chosen while we already had a directional sense of how things were trending. A window logged on day two of a ten-day span, when you already knew what happened on day one, is not a clean prospective window. It is a window with a running start.

The second attempt was more structural. We moved window-setting into the same step as rating publication — the two records had to be created together, or neither could be published. This was the right fix, and it was also the one that forced us to confront how often we had been publishing ratings without having thought seriously about how we would grade them.

Eighteen months of calibration data we could not fully trust

The honest cost was that we had to mark a large portion of our historical calibration figures as unreliable. Not fabricated — the underlying ratings and outcomes were real — but graded under conditions that introduced a systematic upward bias in apparent accuracy. We could not simply regrade them, because the window choices themselves were part of the contamination. Retroactively imposing a fixed window on data that had already been collected under a flexible one does not recover the original signal. It just replaces one kind of distortion with another.

This was the part that stung. The calibration record is, for us, the whole point. A shop that publishes its own error rate and then discovers that the error rate was measured incorrectly is not in an enviable position. We had spent those eighteen months believing we were slightly better calibrated than we actually were. The gap was not enormous — Renata estimated it at roughly six to eight percentage points of apparent accuracy that could not be attributed to genuine performance — but it was consistent, and consistent errors in calibration are the kind that compound.

There is a longer version of this problem we had already documented elsewhere: the specific habit of moving the window to protect a rating we liked. That piece was written as a confession about a single case. What Renata's audit showed was that the single case was not a single case. It was the median behavior.

We also had to sit with the fact that we had written doctrine about this exact failure — and violated it anyway. That is a different kind of embarrassing than simply not knowing better. We knew better. We had published knowing better. We just had not built the system to enforce it.

What the new window protocol actually changed in practice

The structural fix — window and rating published together or not at all — held. It is now the oldest standing rule in our publication process, which is a strange thing to say about something we added after a failure, but that is how most of our rules were made. We did not arrive at them by reasoning from first principles. We arrived at them by doing the wrong thing long enough to notice.

The practical effect was smaller than I expected and more important than I expected, simultaneously. Smaller because the day-to-day experience of rating did not change much — we were already thinking about windows, just not formalizing them early enough. More important because the calibration figures that came out of the new protocol were consistently less flattering, and consistently more stable. Ratings that looked well-calibrated under the old system sometimes looked ordinary under the new one. Ratings that had looked mediocre turned out to be better than we thought. The distribution shifted, and the shift was informative.

One thing we did not fix, and I want to be clear about this: the window length itself is still a judgment call. We have defaults — different spans for different sports, different activity cadences — but the choice of how long a window should be for a given player type is not something the protocol resolves. It is still set by a person, at publication time, with whatever biases that person carries. The protocol prevents us from moving the window after we have seen results. It does not prevent us from choosing a convenient length before we have. That is a problem we are aware of and have not solved.

Instinct plays a role here that I find difficult to quantify honestly. The colleagues who set windows most consistently well are not the ones who follow the protocol most rigidly — they are the ones who have graded enough ratings to have a feel for what a fair window looks like for a given situation. The protocol captures the floor. It does not capture what they are doing above it.

We fixed the mechanism and the calibration numbers improved, and I am still not certain whether the improvement reflects a genuine gain or just a different kind of flattering illusion — one we have not found the audit for yet. The window is fixed in advance now. Whether we are fixing it at the right length is a question the protocol does not answer, and probably cannot.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top