PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Corpus Entry That Was Correct and Useless

The entry had 214 games behind it. The stat was logged consistently, the source was reliable, and the number had barely moved in two seasons — which we took, at the time, as a sign of stability. We ran it through the gates, it cleared every one, and we published a rating we were quietly proud of. The player was real, the games were real, the arithmetic was right. The rating was useless.

The problem wasn't noise. Noise we know how to handle: you wait for it to resolve, you shrink the estimate, you hold the threshold. This was something quieter and harder to catch. The corpus was measuring the right thing in the wrong frame. Every entry was accurate. The whole was meaningless. It took us the better part of a month to understand what had happened, and another two weeks of uncomfortable conversation before we agreed on what to do about it.

I want to write about that, because the version of this failure that ends in a clean lesson is a lie, and the version that ends in "we fixed it and moved on" is almost as dishonest. What we actually have is a process that caught the problem late, a corpus that is now slightly better and still imperfect, and a standing reminder that technical correctness is not the same thing as usefulness — a distinction the corpus has a way of collapsing when you're not watching.

Why Everyday Things Exist

Discover the surprising reasons behind the things, rules, habits, and systems we encounter every day.

Learn more

What "technically correct" actually looked like in this entry

The player was a basketball forward. The stat in question was a counting number — the kind that accumulates across a game and gets logged as a single figure in the box score. Nothing exotic. We had pulled 214 of those figures, spanning two and a half seasons, from a source we use regularly and trust as much as we trust anything.

The issue was context. The figure we were logging was a raw count, and the player's role had shifted twice over those two and a half seasons — once when a teammate returned from injury, once when a coaching change altered how the team ran its offense in the third quarter. The raw count didn't know about any of that. It kept accumulating. We kept logging it. The entry grew, which made it feel more reliable, which made us less likely to question it.

Remi was the one who finally asked the question out loud: are we measuring what we think we're measuring, or are we measuring the sum of three different things that happen to share a name? The answer, once we looked, was the second one. The 214 games contained three distinct player configurations, each with its own baseline, averaged together into a single distribution that didn't accurately represent any of them.

"It passed every gate because the gates check for data, not for whether the data is coherent. We never wrote a gate for that."
— Remi

She was right, and the observation stung a little, because we had written about how much of the real work lives in the gates. The gates were doing their job. The job wasn't big enough.

How we tried to salvage the entry rather than discard it

The first instinct — mine, I'll own it — was to segment the corpus. If the player had operated in three distinct configurations, then split the 214 games into three groups, establish a baseline for each, and use the most recent configuration as the live estimate. The logic was reasonable. We had done something like it before when an entry kept drifting without an obvious cause, and segmentation had helped clarify what was driving the movement.

We tried it. The most recent configuration gave us 31 games. That's not nothing, but it's not 214 either, and the whole reason the entry had felt solid was the depth behind it. Strip out the irrelevant history and you're left with a sample that a single bad stretch could distort significantly. We applied shrinkage, which pulled the estimate toward the base rate for the position, and the resulting number was so conservative it was barely distinguishable from the prior.

The second attempt was to treat the role shifts as covariates — to build a model that adjusted the raw count for context and produced a normalized figure we could actually compare across time. This was more intellectually satisfying and took about three weeks. At the end of those three weeks, Dara ran it against the grading archive and found that the normalized figure was no better calibrated than the conservative shrunk estimate from attempt one, and in two of the five test windows it was worse.

We had spent three weeks building something that performed worse than the thing it was supposed to replace. That's the part of the story I didn't want to write, but here it is.

The cost: a month of ratings built on a distribution that didn't exist

We didn't catch the problem until the fourth week. That means we published ratings during weeks one through three that drew on the original, unexamined corpus entry — the one with 214 games of mixed context averaged into a single number. Those ratings were graded, as everything is. They were not good. The hit rate on that player across those three weeks was low enough that it would have triggered a calibration review under normal circumstances, but it happened to coincide with a stretch where several other entries were also underperforming, and the signal got buried in the noise of a bad overall period.

This is the part of corpus maintenance that nobody talks about honestly: a bad entry doesn't announce itself. It quietly degrades the ratings that depend on it, and if enough other things are also going wrong at the same time — which they have a way of doing — you can miss it for longer than you should. We missed it for a month. The mistake of building around the wrong version of a stat had cost us a similar window the previous year, and we had told ourselves we'd built better detection since then. Apparently not.

There was also a subtler cost. The entry had been in the corpus long enough that other entries referenced it indirectly — players on the same team whose baselines had been partially calibrated against this one. When we corrected the original, we had to re-examine six adjacent entries. Two of them needed adjustment. None of the adjustments were large, but the cascade was a useful reminder that the corpus is not a collection of independent observations. It is a network, and a bad node has neighbors.

What the entry looked like after we stopped trying to fix it and started over

We kept 31 games. Everything before the second role shift was archived — not deleted, because we don't delete, but flagged as context-incompatible and excluded from the live distribution. The entry now reads as thin by our own standards. We require a minimum of 40 games before a corpus entry is considered stable, and 31 doesn't clear that. The player is currently gated out on sample size, which means no rating is published until the current configuration accumulates enough history to say something reliable.

That outcome — a player with two and a half seasons of data who is now ungradeable — is uncomfortable. It looks like we made things worse. I think we made things more honest, which is different, but I also understand why those two things are easy to confuse from the outside.

What we kept from the normalization attempt was not the model itself but the diagnostic it produced. The process of building the covariate adjustment forced us to articulate, precisely, what we believed the role shifts had changed about the player's expected output. That articulation became a written note attached to the entry, which means the next analyst who opens it doesn't have to reconstruct the history from scratch. Whether that note is accurate is a question we won't be able to answer until there's enough data in the current configuration to grade it against. The principle that ungradeable ratings shouldn't be published is one we hold firmly, which means we're waiting.

Dara's summary, delivered at the end of the review, was shorter than mine: "We had a lot of data about a player who doesn't exist anymore." I've been thinking about that framing since. It's not quite right — the player exists, the games happened — but it captures something true about what a corpus entry actually is. It's a model of a person at a particular point in time, and people change configurations, and the corpus has to decide whether to follow them or let the old model go.

We still haven't written a gate that checks for contextual coherence across a corpus entry. We've talked about what it would look like — some kind of structural break detection, a flag that triggers when the rolling baseline shifts by more than a threshold across a defined window. It's on the list. Whether it would have caught this one before week four, I genuinely don't know. The break in this case was gradual enough that even in hindsight I'm not sure where I would have drawn the line.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top