A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

A Corpus Entry We Stopped Trusting

Somewhere around the third week of a hockey stretch, we stopped pulling a particular corpus entry into active scoring. Nobody called a meeting about it. There was no flag in the system, no note in the log. We just quietly routed around it the way you route around a loose step — instinctively, without writing down that the step is loose. It took a calibration review six weeks later to surface what we had been doing, at which point we had to explain to ourselves why a record with 140 game entries and a respectable grading history had been silently demoted to decoration.

The honest answer was that we could not fully say. The entry had not failed a gate. The sample was large enough — we hold 80 games as a minimum threshold before anything goes into active scoring, and this one had nearly double that. The stat type was one we trusted. The player, entirely invented for our corpus the way all of them are, had been consistent across two prior seasons. And yet something in the texture of the recent results had started to feel off to at least two of us, independently, without coordination, and neither of us had written it down.

That is the part that bothered me most. Not the distrust itself — distrust is sometimes correct — but the fact that we had acted on it silently, which meant we had no record to grade. A shop whose entire premise is that it publishes its own error rate had, for six weeks, been quietly running a judgment it could not score.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

The entry that kept passing and kept feeling wrong

The corpus entry in question covered a hockey player — invented, like all of them — tracked primarily on a counting stat that had been stable for two and a half seasons. Stable meaning: the variance was low, the grading window results were within expected bounds, and the line-reading step had not flagged any persistent gap between what the market expected and what the player produced. By every formal measure, the entry was healthy.

The informal measure was harder to describe. Over roughly twelve games, the results had been correct in aggregate but strange in distribution. Not wrong — correct. But the shape of the correctness had changed. Where the entry used to produce results that clustered tightly around the expected value, the recent games showed a pattern of hitting the grade on the far edges: outcomes that were technically within the window but that required the full width of it to qualify. We were right in the way that a clock is right twice a day.

This is a known failure mode. We wrote about it in a piece on entries that are technically correct and useless — records that survive grading without actually telling you anything reliable about a player's performance distribution. The entry here was not quite that, but it was drifting in that direction, and we had not formalized the concern enough to act on it through any documented channel.

Remi, who handles most of the hockey corpus maintenance, put it plainly when we finally sat down with it:

"I knew I didn't trust it. I just couldn't point at a number and say that's the number that broke. It was more like the record had started to feel like a stranger wearing a familiar coat."
That is not a methodology. But it is also not nothing, and the harder question was what to do with a feeling we could not score.

How we tried to find the cause inside the record itself

We ran the standard diagnostics first. Recency weighting: we pulled the last 20 games against the full 140 and looked for a structural break — a point at which the distribution changed shape. There was a soft inflection around game 110, but not clean enough to call a break with any confidence. The kind of inflection that, if you are looking for something, you will find, and if you are not, you will miss entirely.

We checked the stat version. This was a lesson from an earlier failure — we once built an entire corpus strand around the wrong version of a stat and did not catch it for a full season because the wrong version correlated well enough with outcomes to survive grading. Here, the stat version was correct and consistent. No changes to how it had been recorded across the window.

We checked for context contamination — whether the recent games had been played in conditions that differed systematically from the historical baseline. Line combinations, opponent quality proxies, home and away splits. Nothing cohered. The entry looked, on every formal axis we could apply, like a healthy record that had recently been unlucky in a particular way. The distribution shift was real but not large enough to trigger any of our documented thresholds for suspension or review.

What we did not do, and should have, was treat the informal distrust as data and log it explicitly. If two people who work with a corpus every day develop independent reservations about an entry, that is a signal worth recording even when it cannot be quantified. Instead, we acted on it without recording it, which meant we could not grade the decision, which meant we could not learn from it either way.

What the silence in the log actually cost us

The direct cost was small. Six weeks of hockey, one entry, routed around rather than suspended. The ratings that came out of that stretch were not obviously worse for the absence. But the indirect cost was the kind we tend to underweight: we had introduced a gap between what the system said it was doing and what it was actually doing, and gaps like that compound.

When the calibration review surfaced the routing, we had to reconstruct the reasoning from memory. Remi remembered the texture of the concern. I remembered the games that had felt off. Neither of us had written anything down. So the review could not determine whether we had been right to distrust the entry — because we had no record of what we had expected to happen during the six weeks we avoided it, and therefore nothing to grade the decision against. The shop had made a judgment and then made it invisible.

This is related to a problem we have written about from the other direction: the corpus entry we kept updating without fully understanding why the updates felt necessary. In that case, the quiet maintenance was eroding something we had not named. Here, the quiet avoidance was the same mechanism in reverse — action taken below the threshold of documentation, which is the threshold that matters for a shop that grades itself.

The entry, it turned out, probably did have something wrong with it. When we finally suspended it formally and ran a prospective window, the distribution shift persisted and widened. Our instinct had likely been correct. But "likely correct" is not the same as "graded correct," and a shop that only counts the ones it can score has to hold that distinction even when it is uncomfortable.

What the process looks like now for entries we cannot explain

We added a category we did not previously have: provisional suspension with stated reason, where the stated reason is permitted to be "informal concern, not yet quantified." An entry in that category is not used in active scoring, but the suspension is logged with a date, the names of whoever flagged it, and whatever description they can give of the texture of the concern — even if that description is as vague as Remi's coat metaphor. The log entry exists so the decision can be graded later.

The threshold for entering that category is low on purpose. Two independent reservations from people who work regularly with the corpus are enough. We are not requiring a number. We are requiring a record. Those are different things, and conflating them was the error — we had assumed that if a concern could not be quantified it should not be logged, when the correct rule is nearly the opposite: the concerns that resist quantification are exactly the ones most likely to disappear without a trace.

We also now run a short prospective window on anything in provisional suspension — a defined number of games after which we either restore the entry, suspend it formally with documented cause, or close it out of the corpus entirely. The window forces a resolution. Without it, entries tend to stay in the informal purgatory indefinitely, which is its own kind of corruption. There is a version of this problem where a colleague's instinct outscores the formal record not because the instinct is magical but because it is picking up something the corpus has not yet named. We want to be able to grade those cases, and you cannot grade what you have not logged.

The hockey entry was eventually closed. The prospective window confirmed the distribution shift was structural, not random. Whether we would have caught it faster through the formal channel than through the informal routing, I genuinely do not know. The timing might have been identical. What would have been different is the record — and the record is what the shop runs on.

What I keep returning to is the question of whether there is a meaningful difference between a judgment that is correct and a judgment that is graded correct. For most of the work we do, I have always assumed those converge over enough time. This case made me less sure. The entry probably deserved to be sidelined. We sidelined it. But the version of the shop that could not say why it did so, and could not score whether it was right — that version is not doing the thing we say we are doing. Whether that gap is ever fully closeable, or whether some portion of what makes a corpus feel trustworthy will always resist the log, is something I do not have a clean answer to.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top