PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Grading Window Has to Be Fixed in Advance

For about eight months we graded our ratings on a rolling basis, which sounds reasonable until you understand what "rolling" actually meant in practice. It meant that if a player we had rated highly went cold for three weeks, we waited. We told ourselves we were being fair — giving the rating time to resolve. What we were actually doing was leaving the window open until the result looked better, then closing it.

Nobody decided to do this. It accumulated. The window had never been written down anywhere, so each person who ran the grading pass made a quiet local judgment about when enough data had arrived. Riku, who runs most of our calibration checks, noticed it first: our stated confidence on basketball ratings was tracking our observed accuracy suspiciously well — almost too well for a shop that is wrong as often as we are. That is not a compliment. That is a sign the grading is soft somewhere.

When we dug into it, the problem was exactly what you would expect. We were not consciously cheating. We were doing something worse: we were making individually defensible decisions that, in aggregate, systematically favored the ratings we were most confident in. The window was closing later for the confident calls and earlier for the uncertain ones. The calibration looked clean because we had accidentally built a process that cleaned it for us, after the fact, every time.

Business Cash Manager

See what cash is truly available after bills, payroll, taxes, and reserves—before you spend.

Learn more

What an Unfixed Window Actually Measures

A grading window is the span of time between when a rating is issued and when it is scored against what happened. In principle it should be the same for every rating of the same type. In practice, the temptation to adjust it is constant and it arrives disguised as rigor.

The most common disguise is injury. A player we rated highly picks up a minor knock in the first period, plays reduced minutes, and the underlying stat comes in under the line. Did the rating fail? The instinct — and it is not an unreasonable one — is to say the rating was never really tested. So the window gets extended. We will grade it next time he plays a full game. That extension is, on its face, defensible. It is also, in aggregate, a systematic bias: we extend for players we rated confidently, because those are the players we have an argument to protect.

The other disguise is variance. A soccer player we rated as a consistent creator has two poor games in a row. The corpus says this happens — it is within the normal scatter of his historical output. So we wait for the scatter to resolve. Again, defensible in isolation. Again, a thumb on the scale in aggregate. The window is not fixed; it is a negotiation, and the negotiation is always with ourselves.

This is precisely why we have a standing rule — borrowed from a harder lesson — that you write the threshold down before you see the data. The window is a threshold. If it is not written before the rating is issued, it will be written after, and after is too late.

How We Tried to Standardize the Close Date

The first fix was blunt: every rating gets a close date stamped at the moment of issue, calculated mechanically from the type and sport. Basketball player ratings close after the player's next five appearances. Baseball ratings close after ten. Hockey after six. Soccer after eight. Tennis after three matches, because the sample is inherently smaller and we had learned from earlier work — partly through watching what baseball taught us about our hockey gates — that sport-specific thresholds matter more than we had assumed.

The close date goes into the record at issuance and cannot be changed. If the player is injured and misses four of those appearances, the window still closes on the fifth. If the stat in question was genuinely untestable — the player did not play at all — the rating is marked incomplete, not extended. Incomplete is a separate category with its own calibration track. We do not fold incompletes into the hit rate, but we do not hide them either.

Riku pushed for one additional rule that turned out to matter more than the date itself: the grader cannot be the person who issued the rating. This was not about distrust. It was about the ordinary human tendency to remember why a call was reasonable and to let that memory influence how generously you read the outcome. We had been grading ourselves generously for months without knowing it, and the self-grading was the mechanism.

"The problem isn't that people cheat. It's that remembering your own reasoning makes you a bad witness to your own accuracy. You grade the intention, not the result."
— Riku

So the rule became: the issuer flags the close date; someone else runs the grade. On a small team this is awkward — there are weeks when the same two people are doing everything — but even the mild friction of handing it off changes the behavior. The grader who did not issue the rating has no memory to protect.

Where the Fixed Window Made Us Wrong in a New Direction

Fixing the window solved the soft-calibration problem and immediately introduced a harder one: some ratings were now being graded on genuinely uninformative samples. A hockey player rated on defensive contribution who spent his first four appearances killing penalties in lopsided games was graded on five appearances of context-contaminated data. The window closed on time. The grade was clean. The grade was also measuring something adjacent to what the rating had claimed.

We caught this because the calibration numbers for hockey ratings — specifically for defensive players — got worse after we standardized the window, not better. That is the kind of result that makes you sit with it for a while. We had fixed the bias and the accuracy dropped. The most likely explanation was that the old soft window, for all its problems, had occasionally been waiting for genuinely informative samples to accumulate. The fixed window was principled and sometimes useless.

There is a version of this problem that the corpus work surfaces differently: a technically correct entry that tells you nothing about the question you are actually asking. We had written about that kind of entry before in a different context, but we had not applied the same skepticism to our grading samples. A grade can be technically correct — issued on time, closed on time, scored against real outcomes — and still be measuring the wrong thing.

The fix for this is not to reopen the window. It is to be more precise about what a rating is claiming before it is issued, so that the grading sample is defined by the claim rather than by the calendar. We are still working out how to do that without creating a new version of the same motivated reasoning we started with.

What the Fixed Window Forced Us to Keep Honest

We kept the fixed window and the split between issuer and grader. Both are now in the written process, not just in practice. The close date is logged automatically when a rating enters the record, and the log is not editable by the person who issued it. This is a small technical constraint that does the work a policy alone would not.

We also kept the incomplete category, and started treating it as its own signal. A rating type that accumulates a lot of incompletes — players who are unavailable, or whose appearances are too context-contaminated to grade — is telling us something about the gate that preceded it. If players are getting through to a rating and then not being gradable, the gate is probably not doing its job. The doctrine on this connects to something we had already established: you never delete a rating, you only grade it. An incomplete grade is a grade. It goes in the record. It is counted when we report how many ratings resolved cleanly.

The calibration on basketball ratings — the sport where the problem was most visible — took about four months to restabilize after we changed the process. During that period the numbers looked worse than they had before, which was accurate: they had always been worse than they had appeared, and now the appearance matched the reality. That is the only version of improvement the shop knows how to recognize. Not a better number. A more honest one.

What did not change is the underlying belief that sample size beats recency. The fixed window does not override that — it just means the window is long enough to accumulate a real sample and short enough that it closes before we know the answer. Getting both of those right at the same time, for every sport and every stat type, is not a solved problem. It is more of an ongoing argument.

We still occasionally catch ourselves wanting to extend a window — not fraudulently, just with a reason that sounds good. The question I have not answered to my own satisfaction is whether there is a principled way to distinguish "this sample is genuinely uninformative" from "I remember why this rating was reasonable and I would like more time." I am not sure those two things are separable from the inside.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top