A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Gate We Passed Because the Corpus Looked Full

The gate was supposed to catch candidates without enough history behind them. It had a threshold — a minimum number of logged performances before a player could be scored — and it had been working reliably for long enough that we had stopped thinking about what it was actually measuring. That is the part we got wrong. We stopped thinking.

The specific failure happened during a stretch of basketball evaluations late in a season. A candidate cleared the corpus gate comfortably. The entry count was well above threshold. The data was recent enough to satisfy the recency check and the gate moved the candidate forward without a flag. What we did not notice — what the gate could not notice, because we had not built it to — was that a significant portion of those logged performances described a different version of the player. Same name in the database. Different role, different minutes, different physical situation. The corpus looked full. It was full. It was just full of the wrong thing.

We have written before about cases where the corpus entry itself was the problem — entries that survived every gate and still lied tend to do so quietly, which is what makes them hard to catch. This one was louder in retrospect. The signal was there if we had been looking at the right dimension. We were looking at count. Count said fine. Everything downstream said fine too, right up until grading.

Learn to Analyze Data Like a Front Office

A free online course in data analytics with Python: statistics, visualization and finding the signal in the numbers.

Learn more

What "full" meant to the gate and what it should have meant

The corpus gate, as we had written it, asked one question: does this candidate have at least N logged performances in the relevant stat category? We had set N after a calibration exercise a couple of years prior. That exercise showed that ratings built on fewer than N entries had wider error distributions and worse calibration, so we picked the number and enforced it. The logic was sound. The implementation was sound. The problem was that the question was too narrow.

What the gate did not ask was whether those N entries described a coherent player. It did not ask whether the role had changed. It did not ask whether the minutes had shifted by forty percent. It did not ask whether an injury from the prior season had altered the player's physical profile in ways that made the earlier numbers a different sport from the current one. It asked: how many rows? The rows said enough. The gate passed the candidate.

Priya put it plainly when we were writing up the miss:

"We built a gate to check whether the corpus existed. We should have built one to check whether it was coherent. Those are not the same question and we knew that — we just didn't make it structural."

She was right. We had discussed corpus coherence in the abstract before. We had even flagged it as a known gap during an earlier review of the gate architecture. The note existed. We had not acted on it, and the reason we had not acted on it is the reason most known gaps stay open: the gate was passing things correctly often enough that fixing the gap did not feel urgent.

The check we ran after the grade came back wrong

Grading flagged the rating as a miss. The direction of the error was not ambiguous — the candidate performed well outside the range the rating implied, in a way that suggested the underlying expectation was built on a player who no longer existed in that form. We pulled the corpus entry and looked at the row-level data for the first time since the rating had been issued.

The breakdown was visible immediately once we were looking at it. Of the entries that had cleared the count threshold, roughly a third predated a significant role change. Another cluster sat in a period when the player had been returning from injury and operating on restricted minutes. The remaining entries were genuinely current and relevant, but there were not enough of them alone to have cleared the gate. The gate had passed the candidate on the strength of data that was, in a meaningful sense, about someone else.

We ran the rating again using only the coherent subset of the corpus — the entries from the current role, current minutes, post-recovery. The revised estimate was materially different. Not dramatically different in absolute terms, but different enough that the graded miss would have been a much narrower miss, possibly not a miss at all. The corpus had the answer in it. We had just been counting rows instead of reading them.

This is the kind of thing that feels obvious in retrospect and genuinely was not obvious in process. The gate was doing its job. The job was too small.

What the inflated count hid and what we missed by trusting it

The direct cost was a graded miss on a basketball rating. That is what the record shows and we are not going to dress it up. But the more interesting cost was methodological: we had been running this gate for long enough that other ratings had almost certainly passed through it on similarly hollow corpora. We could not go back and regrade them all, but we could look at the distribution of misses over the relevant period and ask whether any of them clustered around candidates with long histories and significant role transitions.

They did. Not dramatically — this was not a systematic collapse, and the calibration numbers for the period were not obviously broken. But there was a mild pattern. Candidates with high entry counts and identifiable inflection points in their history graded slightly worse than candidates with high entry counts and stable histories. We had not noticed this because we had not been looking for it. The gate count was high; we had moved on.

The other cost was to the shrinkage step. Adding more data to a corpus does not automatically improve a rating — we have written about that elsewhere — but we had been treating count as a proxy for reliability when building our confidence intervals. A candidate with more corpus entries got a tighter interval, which meant less shrinkage toward the base rate, which meant more confidence in an estimate that was, in this case, built on a partially fictional player. The inflated count had inflated our confidence. That is a compounding error, and it is the kind we are most susceptible to because it feels like rigor.

The coherence check we added and the thing we did not change

We added a coherence sub-check to the corpus gate. It does not ask whether the corpus is large enough — the count threshold still exists and still does that. It asks whether the corpus is stable enough: specifically, whether the candidate's role and availability profile in the most recent window matches the profile across the full corpus entry. If the divergence exceeds a defined threshold, the candidate gets flagged for manual review rather than passed automatically.

The threshold required calibration of its own. We set it conservatively at first and found we were flagging too many candidates for review — players go through minor fluctuations in role all the time, and not all of them represent the kind of structural change that invalidates historical data. We widened the threshold, which meant we were probably still passing some candidates we should flag. That tradeoff is unresolved. The gate is better than it was. It is not correct.

What we did not change was the base count threshold. There was a brief discussion about whether the miss meant the number itself was wrong — whether we should raise it, or add a separate threshold for the coherent subset. We decided against it, partly because the calibration exercise that had set the original number was still the best evidence we had, and partly because adding gates in response to individual misses is its own failure mode. A gate added in reaction to a specific error tends to be too narrow, too brittle, and tends to create a new gap somewhere adjacent. We have done that before. The record on it is not good.

The coherence check is in the gate architecture now. It has flagged candidates that would have passed cleanly under the old version. Some of those flags led to manual reviews that changed the rating; some led to reviews that confirmed the original estimate. We do not know yet whether the check is well-calibrated — it has not been running long enough to say. That is the honest answer, and it is the one we are going with.

The corpus gate was built to ask whether enough data existed. The question we actually needed it to ask was whether the data was still describing the same thing. Those feel like the same question until they are not, and I am not entirely sure how you build a gate that knows the difference before a miss tells you to look.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top