PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Chasing a Pattern That Was Two Data Points

Somewhere around the third week of a basketball stretch run, Dara flagged something she called "a clean directional signal." A forward — invented here, as always — had cleared a particular counting-stat threshold in back-to-back road games against high-pace opponents. The corpus showed the threshold was meaningful over full seasons. The gate that was supposed to ask how many times has this actually happened had been set to a minimum of two. Two. We had written the gate ourselves, at some point, apparently in a generous mood, and we had never gone back to question it.

The rating survived every step of the process. It cleared availability, cleared the recency filter, cleared the line-reading stage where we ask what the market's expectation implies about underlying belief. It was capped, shrunk toward the base rate, and graded on schedule. The grading is what embarrassed us. Not because the player performed badly — he performed fine, roughly at his seasonal average — but because we had called the pattern structural when it was, in the most literal sense, two data points dressed up in the language of a trend.

We write up the misses because a shop whose entire premise is honest self-grading has no defensible reason to quietly move on. This is one of the worse ones, not because the error was exotic but because it was so ordinary. We had the rule. We had just set the rule too low, and then stopped reading our own work.

Modern Test Data Engineering

Practical guides for generating, managing, and validating test data across modern systems.

Learn more

The Gate That Let Two Observations Pose as a Pattern

The gate in question was designed to enforce a minimum occurrence count before any situational split — road game, high-pace opponent, back-to-back — could be treated as evidence of a real tendency rather than noise. The logic was sound. A player who has exceeded a threshold in thirty road games against fast teams is saying something. A player who has done it twice is saying almost nothing, and the gap between those two statements is the entire reason the gate exists.

What we had done, at some earlier point that nobody could precisely date, was set the minimum to two. The rationale, probably, was that we wanted the gate to catch true single-game flukes while still allowing the process to run on players with limited recent road exposure. That is a reasonable concern. It is also the kind of reasoning that sounds careful and produces a gate with no real teeth.

The situational corpus for this player had exactly two qualifying games. Both went over the threshold. The gate read "two of two" and passed him through. At no point did the system pause to note that "two of two" and "thirty of forty-seven" are not the same kind of statement, even though we have written about the damage a small sample does to an otherwise disciplined process. We knew this. The gate just didn't.

How the Rating Got Built Anyway, Step by Step

Once the gate passed him, the rest of the process ran normally, which is part of what makes this instructive. The line-reading stage noted that the market's expectation for this player on this night sat slightly below his two-game situational average. We treated that as mild evidence of disagreement between the situational split and the market's implied belief — which is exactly what you are supposed to do when sources diverge, and exactly what you should not do when one of the sources is two games old.

The shrinkage step pulled the estimate back toward the base rate, as it always does. The base rate for this player across his full corpus was unremarkable — he was a consistent mid-tier performer, nothing more. Shrinkage should have done more work here. Because the situational split was so thin, the weight assigned to it in the blended estimate should have been close to zero. Instead, the formula gave it a weight proportional to its apparent consistency: two for two is a 100% hit rate, and the formula, reading that number without any adjustment for sample fragility, treated it as meaningful signal.

"The formula saw 1.000 and got excited. It does not know that 1.000 on two attempts is one lucky bounce from 0.500." — Dara

She was right, and she said it after the fact, which is the only time any of us tend to say the obvious things. The rating was published, internally, with a confidence tier that now reads as quietly absurd. We had written the threshold down before seeing the outcome — that part we did correctly — which is the only reason we have a clean record of exactly how wrong the confidence was.

What the Grading Window Showed, and What We Had Told Ourselves

The player's actual performance that night sat at his seasonal average. Not a collapse, not a breakout — just the number a full-corpus reading would have predicted if we had trusted the full corpus and ignored the two-game split entirely. The grading log records the rating as a miss on confidence: we had stated a higher certainty than the outcome frequency justified, which is the specific failure mode calibration exists to catch.

The deeper cost was not that one rating. It was that the two-game gate had been running for longer than we realized. When we audited the prior three months of output, we found four other instances where a situational split had cleared the gate on fewer than five qualifying observations. Two of those had graded well, which is the part that stings: the ones that grade well after a thin-sample pass are the ones that teach the formula the wrong lesson. We had been rewarding the same structural error in two directions — congratulating the process when a lucky thin-sample rating graded out, and treating the misses as isolated bad luck.

Marcus looked at the calibration log for those four cases and noted that our stated confidence on all of them averaged about twelve percentage points above our observed accuracy. Twelve points is not catastrophic, but it is consistent, and consistency is what turns a small miscalibration into a systematic one. The error was not random noise. It had a shape, and the shape pointed directly at the gate minimum.

What we had told ourselves, implicitly, was that the shrinkage step would catch anything the gate missed. It does not work that way. Shrinkage pulls an estimate toward the base rate, but it does so in proportion to the estimate's apparent uncertainty. An estimate built on a 100% hit rate, however small the sample, presents as certain, and shrinkage treats it accordingly. The gate is the last line of defense against thin situational splits. Ours was set to a number that offered almost no defense at all.

The Gate Minimum We Changed, and the One We Are Still Arguing About

The immediate fix was straightforward: the minimum occurrence count for any situational split to receive positive weight in the blended estimate went from two to eight. Eight is not a magic number. It is the point at which, in our corpus, the variance on a split's hit rate drops to a level where shrinkage can do meaningful work with it. Below eight, the formula's confidence weighting produces estimates that are too far from the base rate to be honest, and the grading record now confirms that.

We also added a visible flag to any rating where the primary situational driver has fewer than fifteen qualifying observations. Not a disqualification — a flag. The distinction matters because sometimes a player genuinely has limited exposure to a specific situation, and "we cannot say much" is itself useful information. The flag makes the thinness visible rather than hiding it inside a blended number that looks authoritative.

What we are still arguing about is whether the minimum should be higher. Dara thinks eight is still too low for pace-of-play splits specifically, because pace varies enough year to year that a player's games against fast opponents from two seasons ago may not describe anything about tonight. Marcus disagrees, on the grounds that reducing the qualifying pool too aggressively will leave us with nothing to say about most players in most situations, and silence is not obviously better than a well-flagged thin estimate.

He is probably right that silence has its own costs. The part I keep returning to is that we published the confidence tier without anyone in the room asking how many games the split was built on. That question should be automatic. We have written about the discipline of not publishing what cannot be graded — and the corollary, which we apparently needed to learn separately, is not publishing what cannot be trusted just because it technically can be graded.

I am still not sure whether the right response to a two-game situational split is a higher gate or just a standing habit of asking the question out loud before the rating moves. Maybe the question is the gate. Or maybe a question that depends on someone remembering to ask it is not a gate at all.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top