A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

The Gate We Added After the Fact and Called Foresight

The miss happened in a basketball corpus entry during what we later called the rotation week — a stretch in late winter when a handful of teams made quiet adjustments to their lineups without any formal announcement. We had a candidate who cleared every gate cleanly. Sample size was adequate, recent activity was present, no availability flag had been tripped. We rated him. He played eleven minutes. The rating was not designed for eleven minutes.

The embarrassing part was not the miss itself. Misses are in the budget. The embarrassing part was what happened in the three days after it, when we sat around the table reconstructing the episode and gradually convinced ourselves that there had been a signal we had failed to act on — a pattern in the candidate's recent minute totals that, in retrospect, looked like a taper. We designed a gate around it. We wrote it into the process document. We gave it a name. And for a while, we referred to it in internal notes as though it had always been there and we had simply forgotten to formalize it.

It had not always been there. We had invented it after the fact and dressed it in the language of foresight. It took the calibration log another six weeks to tell us what we should have known immediately: the gate was not catching anything it was designed to catch. It was catching the memory of one bad night and calling that a rule.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

How a Taper Pattern Became a Gate It Had No Right to Be

The pattern we identified was real, in the narrow sense that it existed in the data. Over the four weeks preceding the rotation week, the candidate's minutes had declined in a shape that looked, if you squinted at it, like a deliberate reduction. Twelve games. Each one slightly shorter than the one before. We drew a line through the points and the line went down.

What we did not ask — and this is the part that still bothers me — was whether that shape would have been visible before the miss, or only after it. When you already know the outcome, a declining trend feels like a warning. When you do not know the outcome, eleven of those twelve games were within a normal range of variance for a rotation player, and the twelfth was the one where the coach made a change we had no way to anticipate. The line we drew was a line through noise plus one genuine event, and we had weighted the genuine event at roughly ten times its proper share.

Rowan, who runs the corpus maintenance side of the shop, said something at the time that I wrote down and have not thrown away:

"We're not building a gate. We're building a monument to the last thing that surprised us."
I thought he was being uncharitable. He was being precise.

What We Built and How We Justified It to Ourselves

The gate, as we formalized it, worked like this: if a player's per-game minutes showed a negative slope over the trailing twelve appearances, and if that slope exceeded a threshold we set at minus-0.4 minutes per game, the candidate was flagged for manual review rather than passed automatically. It was not a hard disqualification. It was a soft stop — a place where a human was supposed to look more carefully before the rating proceeded.

On paper, that sounds reasonable. Soft stops are not inherently bad. The gate that leaked for eight days was a hard stop that should have been a soft one, and we paid for that rigidity. The problem here was not the architecture. It was the premise. We had designed the gate to catch a specific type of event — a coach quietly reducing a player's role — but we had calibrated the threshold against exactly one confirmed example of that event, and that example was the miss that had prompted the gate's creation. We were, in the language of the calibration log, fitting to our own residual.

The justification we used internally was that the gate would have caught the original miss if it had existed. This is true and also completely useless as a design criterion. A gate designed to catch the thing that already happened will always catch the thing that already happened. The question is whether it catches the next version of that thing without also stopping candidates it has no business stopping. We did not ask that question with any rigor. We asked it briefly, answered it optimistically, and moved on.

Six Weeks of Soft Stops That Were Mostly Just Noise

Over the following six weeks, the gate flagged nineteen candidates for manual review. Of those nineteen, three had genuine rotation concerns that we would likely have caught through other means anyway. The remaining sixteen were rotation players whose minute totals varied normally across a twelve-game window and whose subsequent performances were unremarkable. The gate had not identified anything. It had generated work.

The calibration log was more specific and more damning. Our stated confidence on ratings that passed through the manual review step was, on average, slightly lower than on ratings that cleared without it — which is what you would expect, since the review step introduced hesitation. But the observed accuracy on those ratings was indistinguishable from the baseline. The hesitation was not improving the output. It was just making us feel more careful while producing the same results.

This is a version of a problem we have written about before: the difference between a gate that actually filters and a gate that performs filtering. The gate we closed correctly for the wrong reason was a different shape of the same failure — a mechanism that produced the right outcome through a logic that did not hold. Here the logic did not hold and the outcome was not even right. We had the worst of both versions.

What it cost, in the end, was not accuracy. It was clarity about what the process was actually doing. Once a gate exists, it accumulates institutional weight. People stop questioning whether it belongs. We nearly added a second gate in the same vein — a slope-of-slope measure, a second derivative of the taper — before Rowan pulled the document and asked how many confirmed examples we had for the new threshold. The answer was zero. We had been about to fit a gate to a hypothesis about a gate that was already fitted to a single data point.

What Stayed After We Removed the Gate We Should Not Have Built

We removed the minute-slope gate at the end of that calibration window. The process document went back to its prior form, with one addition: a note in the review protocol asking explicitly whether any gate added in the previous ninety days had been introduced in response to a specific miss. If the answer was yes, the gate needed a prospective sample before it was treated as permanent — at minimum, twenty candidates evaluated under the new rule, with outcomes recorded, before the rule was considered validated.

That protocol has since caught two other instances of the same pattern. Neither was as elaborate as the minute-slope gate. Both were softer — small adjustments to thresholds, the kind that are easy to rationalize as refinements rather than reactions. The protocol made them visible. We removed one and kept one, and the one we kept had enough prospective data behind it by the time we reviewed it that we felt reasonably confident it was doing something real. We were careful not to say very confident, because some gates take more than one pass before you understand what they are actually measuring.

The original rotation-week miss stayed in the grading record as what it was: a candidate who cleared every gate that existed at the time, played eleven minutes, and produced a rating that was wrong. Not wrong because the process failed. Wrong because the process did not have the information it would have needed to be right, and no amount of retroactive gate-building was going to change that. Sometimes the corpus is complete and the gates are sound and the outcome is still bad. The entry that survived every gate and still lied is a different flavor of the same truth: a clean process is not a guarantee, and mistaking a clean process for a guarantee is how you end up building monuments to your own surprises.

The question I keep returning to is whether there is a principled way to distinguish between a gate that belongs in the process and a gate that is just grief wearing a methodology costume. The ninety-day protocol helps, but it is not a complete answer — a gate fitted to a single miss can still accumulate a plausible prospective sample if the conditions that produced the original miss happen to recur. I do not have a clean test for this. I am not sure one exists.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top