The Gate We Added Because the Data Scared Us
There is a specific kind of unease that arrives when a corpus returns results that are technically fine and feel completely wrong. The numbers clear every threshold. The candidate has enough sample. The line reads as a credible statement of expectation. And yet something in the distribution is sitting at an angle, and you keep scrolling back through it the way you check a locked door twice. That happened to us in a stretch of basketball evaluations where three candidates in the same ten-day window posted performances so far outside their established ranges that our grading file looked like it belonged to a different shop. Nothing was broken. The method had just met something it had not been built to handle.
The performances were not fraudulent. The players had not been misidentified. The corpus entries were accurate. What we were looking at was genuine volatility — the kind that is real, documented, and completely unhelpful if you are trying to say something reliable about what a player is expected to do on a given night. A player whose output ranges from six to thirty-eight in the same stat, across a legitimate sample, is not a player the method was designed to evaluate. He is a player the method was designed to skip. We had no skip.
So we built one. This is the account of that gate — why we added it, how we designed it, what it caught, and what it cost us that we did not fully understand until the calibration review three months later.
A free online course in data analytics with Python: statistics, visualization and finding the signal in the numbers.
What the Distribution Was Actually Telling Us
The three candidates that triggered this were not outliers in any simple sense. Their career-level means were reasonable. Their sample sizes cleared our minimum. They had passed every gate we had — availability, expected participation, recency, event status. The problem was not scarcity of data. It was the shape of the data.
When we pulled the full distributions, all three showed what a colleague named Dara called, with some understatement, "a very wide personality." Standard deviations running at sixty to seventy percent of the mean. Bimodal-ish histograms with a cluster near the floor and a second cluster near the ceiling and not much in between. These were players who did not have a typical night. They had nights.
"A mean doesn't mean much when the distribution has two modes and neither of them is the mean," Dara said, which is the kind of sentence that sounds obvious until you have spent an afternoon defending a rating built entirely on the mean.
Our corpus was recording what happened accurately. What it was not doing was flagging that the recorded history was structurally unsuitable for the kind of point-estimate rating we were producing. We were collapsing a bimodal distribution into a single number and then treating that number as though it described something. It described the average of two very different players who happened to share a body.
The line, in those cases, was reading the same distribution we were — and in at least two of the three cases, the line looked to us like it was set near the trough, which meant the high-output mode was doing most of the work in our estimate. We were not wrong about the history. We were wrong about what the history was evidence of. That distinction took us longer to articulate than it should have.
How We Designed the Volatility Gate
The gate we settled on used a coefficient of variation threshold — standard deviation divided by mean — with a secondary check on interquartile range relative to median. If a candidate's CV exceeded a fixed threshold and the IQR-to-median ratio confirmed the spread was not just a few extreme events pulling the tails, the candidate was disqualified from rating entirely. Not penalized. Not flagged for human review. Removed.
We set the CV threshold at 0.55 after a calibration exercise on historical data, running the gate backward across roughly fourteen months of basketball evaluations to see what it would have caught. It would have removed those three candidates. It would also have removed a handful of others whose ratings had graded out fine — which was the first sign that we were not only solving the problem we thought we had.
The secondary IQR check was Milo's suggestion. His argument was that a high CV caused by two or three genuinely anomalous games was a different problem from a high CV caused by a structurally bimodal output pattern, and that the gate should only trigger on the latter. He was right about that distinction. Whether the threshold we landed on actually separated them reliably is a different question, and the answer is: imperfectly.
We ran the gate across all five sports in scope. This was, in retrospect, a version of the same error documented elsewhere in our gates history — building something for one context and assuming it travels. The CV threshold that made sense for basketball, where a player's role can shift dramatically within a game, was too aggressive for baseball, where volatility in counting stats reflects something structurally different. We did not catch that for six weeks.
What the Gate Removed That We Did Not Mean to Remove
The calibration review three months out is where we found it. The gate had been running cleanly in basketball. In baseball it had been quietly disqualifying a category of player — high-variance by nature of their role, not by instability — whose ratings had historically been among our better-calibrated outputs. We had removed them because their distributions looked wide. Their distributions looked wide because their role genuinely produced wide distributions, and our prior ratings had accounted for that. The gate had not.
The net effect was that the gate improved our basketball calibration slightly and degraded our baseball calibration measurably. We had solved a problem in one sport by introducing a different problem in another, and because we were watching basketball closely and baseball less so, we did not see it for six weeks. By then we had three months of baseball ratings that were less reliable than the ones we had been producing before we tried to make things better.
There is a version of this story where we caught it at design time. Dara had flagged that the threshold felt sport-agnostic in a way that should have made us nervous, and we had noted the concern and moved on because the backward calibration on basketball looked good. Calling a gate well-designed because it handled the incident that prompted it is a failure mode we have documented before. We repeated it.
We also found, in the same review, that the gate was interacting with our sample-size minimum in a way that produced an unintended selection effect: candidates with shorter histories were less likely to show high CV scores simply because there were fewer games to reveal the spread. So the gate was, in some cases, passing volatile players with thin corpora while blocking stable players with large ones whose legitimate variance had finally had room to show up. The gate was punishing depth. That is a fairly embarrassing property for a shop whose oldest belief is that sample size beats recency.
What the Gate Looks Like Now, and What We Still Do Not Know
We split the gate by sport. Basketball retained a CV threshold, lowered slightly and paired with a minimum sample requirement before it can trigger — a candidate needs at least forty games in the corpus before volatility disqualification is even on the table. The logic is that a high CV across twelve games might just be a short history; a high CV across fifty games is a structural statement about the player.
Baseball got a different threshold, higher, and the IQR check was replaced with a role-consistency flag that Milo built over a weekend and that we have not yet fully stress-tested. Hockey and soccer got sport-specific versions after we found, in the course of fixing baseball, that we had never properly examined whether the basketball parameters applied there either. Tennis, where the relevant stats are different enough that CV behaves oddly, got a separate volatility measure entirely that we are still calibrating.
What we kept from the original design is the principle: a candidate whose history cannot support a point estimate should not receive one. That part was right. The error was in assuming that a single mathematical criterion could identify "cannot support a point estimate" uniformly across sports that have different structural reasons for variance. Forgetting to revisit a gate after its original context has changed is a recurring problem for us, and this gate is now on a six-month review schedule specifically because of that history.
The gate runs. It removes candidates. We believe the removals are mostly correct, in the sports where we have calibrated it properly, and we are genuinely uncertain about tennis. The calibration review will tell us more. It usually does, and not always what we were hoping to hear.
What I keep returning to is Dara's original observation — that a mean doesn't mean much when neither mode is the mean — and whether the right response was a gate at all, or whether it was a different kind of rating that could hold a distribution's shape rather than collapsing it. We chose the gate because it was faster and because the data scared us into wanting a clean removal rather than a harder problem. I am not sure that was wrong. I am not sure it was right either.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.