PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Corpus Built on the Wrong Stat Version

For about fourteen months, we maintained a basketball corpus built around a version of a shot-creation metric that we had quietly redefined at some point in the previous cycle. We did not notice we had redefined it. That is the whole story, really, but the details are worth writing down because the mechanism was subtle enough that I want to make sure we can recognize it again.

The stat in question measured how often a player generated a field-goal attempt — for themselves or a teammate — from a specific zone of the floor. We had been collecting it for two seasons before we decided, during a routine corpus audit, that the zone boundary was too generous. We tightened it. We updated our collection scripts. We did not rename the column. We did not flag the historical data as collected under the old definition. The column was called creation_rate in the old data and creation_rate in the new data, and for fourteen months we compared numbers that were not the same number.

The ratings we produced during that window looked fine. Calibration held within acceptable range for the first two quarters. It was only when we tried to extend the corpus backward — to test whether a proven performer's long record was outperforming shorter hot streaks in our grading — that the seam appeared. The historical numbers and the current numbers disagreed in a way that the players' actual performance did not explain.

Wanderlust y Couture with LuxeSofia

Discover luxury hotels, chic city stays, and beautiful escapes around the world.

Learn more

How Two Definitions Lived in the Same Column for Over a Year

The zone boundary change was not trivial. The old definition included attempts generated from a mid-range arc that extended roughly two feet beyond what we later decided was analytically meaningful. Players who worked that area — a specific type of pull-up shooter who initiates from the elbow — had their creation numbers inflated under the old definition. Under the new one, they looked average. Under the blended corpus we had accidentally built, they looked inconsistent.

Inconsistency, in our method, is penalized. A player who looks inconsistent gets shrunk harder toward the base rate during the cap-and-shrink step. So these players were being pulled down not because their actual performance was volatile, but because we had changed the ruler mid-measurement. We were punishing them for a clerical error and calling it rigor.

Mira was the one who found it. She had been doing a routine check on why a particular cluster of players kept failing the consistency gate despite having what looked like clean underlying performance. She pulled the raw creation numbers across three seasons and noticed that the variance pattern had a hard edge at the exact date we had updated the collection scripts. Before that date, the numbers were higher and tighter. After it, lower and wider. The players had not changed. The definition had.

"The variance wasn't in the players. It was in us. The column name was the same so nobody thought to check what the column meant."

She was right, and I had signed off on the script update without documenting the definition change in the corpus log. That part is on me specifically.

The Restatement We Tried Before Admitting We Needed to Start Over

The first thing we tried was a correction factor. If we knew the old zone was approximately two feet more generous, and we had enough games where both definitions could be computed in parallel, we could estimate a conversion multiplier and apply it backward. This is a reasonable approach when the measurement change is stable and the affected population is uniform. Neither condition held here.

The multiplier varied by player type. The elbow pull-up shooters needed a different factor than the transition creators, who needed a different factor than the off-screen receivers. We had not tagged players by type in the corpus — another gap — so we were trying to apply a single correction to a heterogeneous population. The corrected numbers were better than the uncorrected ones, but "better than a known error" is not the same as "correct," and we knew it.

We ran the corrected corpus for six weeks. Calibration improved slightly. We wrote that up internally as a success, which was the second mistake. The first mistake was the undocumented definition change. The second was treating marginal improvement as validation. This is exactly the dynamic we described when we wrote about a week where everything was confidently wrong — the confidence was the problem, not the data.

After six weeks we pulled the corrected corpus and went back to raw collection under the new definition only, accepting that we had lost the historical depth for this particular metric. The proven-performer advantage that the corpus is supposed to provide — the whole reason sample size beats recency — was gone for creation rate. We were back to shallow samples, which is the thing we trust least.

Fourteen Months of Ratings That Were Grading the Wrong Thing

The honest accounting: every creation-rate-weighted rating we produced between the script update and Mira's audit was grading players against a blended definition that matched neither the old market understanding nor the new one. We do not know exactly how many ratings were materially affected. Our estimate, based on how heavily creation rate was weighted in the basketball model at the time, is that it touched somewhere between a third and half of the basketball output in that window.

The calibration numbers did not scream. That is the part that bothers me most. We talk about calibration as the thing that catches errors, and here it did not catch this one quickly enough. The reason, I think, is that the error was symmetric enough in aggregate — some players were being pulled down, others were being rated more cleanly — that the overall hit rate did not collapse. The error was structural but the signal looked noisy rather than broken. A stat that measures nothing and a stat measured inconsistently can look identical from the outside of the calibration report.

We also did not catch it at the gates. The availability and activity checks are not designed to catch definition drift — they are designed to catch missing data and inactive players. A player with a full season of blended creation-rate numbers passes every gate cleanly. The gates, in this case, were not the right instrument for the failure mode. That is worth noting separately, because some of our gate rules were already known to be imperfect, and this was a different kind of imperfection than the ones we had already catalogued.

What it cost, in practical terms, was the historical depth of the creation-rate series and roughly fourteen months of basketball ratings that we can no longer stand behind with the confidence we assigned them at the time. We have not deleted them. They sit in the grading log marked with a flag that says the underlying metric definition was inconsistent across the window. That flag is not satisfying, but it is accurate.

What the Corpus Log Now Requires When a Definition Changes

We added one rule to the corpus maintenance protocol: any change to a metric's collection logic requires a new column name, a deprecation notice on the old column, and a dated entry in the corpus log that describes what changed and why. The old column stays in the database as read-only. The new column starts accumulating from the change date. They are never merged without an explicit, documented decision about how to handle the seam.

This sounds obvious. It is obvious. We did not do it because the change felt small at the time — a zone boundary adjustment, not a new stat — and because the column name was already established in a dozen downstream scripts. Renaming it felt like more work than the change warranted. That calculation was wrong.

We also added a variance-pattern check to the corpus audit. Once per cycle, for every metric in the basketball and hockey models, we plot rolling variance over time and look for hard edges — places where the variance changes shape in a way that does not correspond to any known change in the player population. A hard edge that coincides with a script update is a warning. A hard edge with no corresponding log entry is an incident.

What we did not do — and I want to be clear about this — is rebuild the historical creation-rate series using the new definition. We considered it. The data exists to do it for most of the affected seasons. But reconstructing a corpus retroactively, with the knowledge of what the ratings produced, introduces a different kind of contamination: you know what you want the history to look like. We decided the flagged gap was cleaner than a reconstruction we could not fully trust ourselves to do without bias. That may be the wrong call. We have not resolved it.

The thing I keep returning to is that the error was invisible precisely because the column name stayed the same. Names are load-bearing in a corpus. They carry the assumption that the thing being measured is the same thing it was last time. When that assumption breaks quietly, the corpus keeps working — it just works on a question you stopped asking without noticing. I do not know a clean way to guard against that entirely, and I am not sure the variance-pattern check we added is sufficient. It caught this class of error after the fact. Whether it would catch a subtler one, I genuinely cannot say.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top