The Corpus Entry That Survived Every Gate and Lied
The gates are supposed to catch the bad ones. That is their entire purpose — disqualify the thin samples, the unavailable players, the entries where the stat being tracked changed definition halfway through a season. We built them carefully, added to them over three years, and by last winter we had eleven of them running in sequence. Most candidates die at gate two or three. The ones that reach gate nine are, in theory, the entries we can actually say something about.
This particular entry reached gate eleven. It had four seasons of performance data across two sports — the player had competed professionally in both, which is uncommon enough that we flagged it as interesting rather than suspicious. The sample was large. The recency was balanced. Every automated check returned clean. We scored it, ran it through shrinkage, and published a rating. The rating was wrong in a way that took us eight months to understand.
What follows is not a story about a gate that failed. The gates worked exactly as designed. It is a story about what gates cannot see, and about the particular way a corpus entry can be immaculate in structure and dishonest in substance at the same time.
Securely manage keys for 60+ AI providers in one encrypted vault instead of juggling them across apps.
What the Entry Looked Like Before We Understood What It Was
The player — invented here, as all our examples are — had logged roughly 180 tracked performances across basketball and soccer over four seasons. The basketball record was the older and larger of the two. The soccer record was more recent, accumulated during what the corpus treated as a clean transition: same player, different sport, same measurement framework applied uniformly.
That last assumption was where the trouble lived, though we did not know it yet.
Our framework tracks output against expectation — what a line implied a player would do, versus what they actually did. The theory is that if you do this across enough performances, you learn something stable about the player: whether they tend to outperform expectation, underperform it, or cluster tightly around it. That clustering is what a gem rating is. It is not a prediction. It is a historical characterization of how a player has related to what the market thought of them.
For this entry, the basketball record showed tight clustering with a slight positive lean. Reliable. The soccer record showed a different pattern — wider variance, less consistent. When we combined them into a single cross-sport rating, the basketball record's size dominated, and the combined rating looked more stable than it should have. We were, without knowing it, technically correct and useless — the rating accurately described the basketball history and said almost nothing true about soccer performance.
How We Tried to Make Two Sports Into One Record
The cross-sport problem is not new to us. We have written before about gates designed for one sport that we stretched across four, and the results were instructive in the same uncomfortable way. The instinct to unify is strong: if the underlying measurement is "how does this player relate to expectation," then the sport should be a context variable, not a structural barrier. We believed that for a long time.
What we tried here was a weighted combination. Basketball performances were weighted by their distance in time — older ones discounted, recent ones amplified — and soccer performances were treated as a separate cluster that would gradually earn more weight as the sample grew. The two clusters were then merged using a blending coefficient we had calibrated on a different set of cross-sport players from an earlier season.
Rowan, who built the original blending logic, was skeptical from the start. Her note at the time, which I have kept:
"The coefficient was calibrated on players who moved between sports with similar physical demands and similar market depth. This one moved between sports with different market depths. The lines in soccer were set by a thinner market. We're treating them as equivalent evidence. They're not."
We ran the blending anyway. The combined entry passed every gate. The gates checked sample size, recency balance, stat-definition consistency within each sport, and availability signals. None of them checked market depth as a variable in evidence quality. That was the gap.
Eight Months of Ratings Built on a Stable-Looking Lie
The rating we published in late winter held through spring and into summer. We graded it on schedule, inside the fixed window, and for the first three months the grading looked acceptable. The basketball performances — which continued to make up the majority of the player's activity — tracked reasonably. The soccer performances did not, but because they were the smaller portion of the record, the overall grade stayed within a range we did not flag.
This is the specific failure mode I find most embarrassing in retrospect: the grading process was honest, and it still let us down. We were grading the combined rating against combined performance, which meant the large, well-behaved basketball sample was absorbing the noise from the small, poorly-calibrated soccer sample. The rating was not lying to us loudly. It was lying to us quietly, in a register the grading system was not designed to hear.
By late summer, the soccer performances had accumulated enough that the pattern became visible. The variance was not random. It was directional — the player was consistently outperforming soccer lines in a way that suggested the lines were being set by a market that had not yet caught up to the player's actual soccer output. Our rating, anchored to the basketball history, was not capturing that. We had built a record that kept getting updated without the updates actually changing what the rating believed.
Eight months. That is how long a structurally clean entry can produce ratings that are systematically miscalibrated before the grading window catches it, if the miscalibration is sport-specific and the sport in question is the minority portion of the record. I do not have a good answer for how to shorten that lag.
What Survived: Separate Ledgers, Slower Merges, Rowan's Coefficient
We split the entry. Basketball lives in its own record now, scored against basketball lines, graded against basketball outcomes, shrunk toward the basketball base rate. Soccer lives separately, scored against soccer lines, graded against soccer outcomes, shrunk toward the soccer base rate. The two records are not merged. If a cross-sport rating is ever needed, it will be computed fresh at the time it is needed, with explicit documentation of the blending assumptions and their calibration source.
That last part — explicit documentation of blending assumptions — is the piece we had skipped before. The coefficient Rowan built was well-reasoned, but it lived in a comment block in a script rather than in the corpus record itself. When we used it on a new case, we did not revisit whether the calibration still applied. We just used it. That is the kind of invisible inheritance that makes a corpus drift without anyone deciding to drift it.
We also added a gate. Gate twelve now checks market depth as a proxy for evidence quality: if the lines in a given sport were being set by a market with fewer than a threshold number of contributing sources during the period in question, the performances from that period are flagged as lower-confidence evidence. They are not discarded — discarding data is almost always wrong — but they are weighted accordingly, and the flag travels with the record through every subsequent step.
What we did not do is assume the problem was solved. The new gate catches the version of this failure we already saw. It probably does not catch the next version. That is the honest thing to say about a gate built in response to a specific miss: it is a scar, not a shield.
Rowan asked, after we closed the entry and split the record, whether the real issue was that we trusted the gates too much — whether passing eleven checks had made us less curious about the entry rather than more. I have been thinking about that since. A long clean run through a process is either a sign that the entry is genuinely solid, or a sign that the process has not yet encountered the particular way this entry is wrong. We cannot always tell which one it is until later, and sometimes not even then.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.