What a Good Month Nearly Cost Us
The month ended and the hit rate was the best we had logged in about two years. Not by a little — by enough that Renata printed the summary and taped it to the whiteboard without saying anything, which is her version of a standing ovation. We had run the grading window, compared stated confidence to observed accuracy, and the gap was the smallest it had been since we rebuilt the corpus after the hockey expansion problem. It felt, briefly, like we had figured something out.
We had not figured anything out. What we had done, without noticing, was spend four weeks making the method progressively more accommodating to the conditions that were producing good results — and calling that accommodation "refinement." By the time the next grading window closed, the calibration gap had widened to a point that erased most of what the good month had suggested. The accuracy number stayed respectable. The calibration number was embarrassing. As we have written before, accuracy and calibration are different things, and confusing them is the shop's oldest recurring mistake.
This is the write-up of that month, and the one that followed it. The Misses desk exists because a shop that publishes its error rate has to show the errors in full, not just report the aggregate. This one is harder to write than most because the error was not a miscalculation. It was a decision that felt reasonable at the time and turned out to be a form of self-congratulation dressed in methodological language.
Victor Draemont’s notes on discipline, judgment, power, and playing the long game.
The Month the Method Started Agreeing With Itself Too Easily
The good results began in basketball. We had a stretch of players whose corpus was deep — three-plus seasons of comparable performance data — and the lines we were reading as evidence about market beliefs were landing close to where our own estimates sat. Not identical, but close enough that the gates were clearing candidates quickly and the shrinkage adjustments were small. The whole process felt frictionless, and that is precisely when the shop should have slowed down.
Instead, we did what apparently comes naturally after a run of good results: we started treating the method's outputs as confirmation of the method's soundness, rather than as outputs that still needed to be checked against the method's own rules. The caps were applied, technically. The shrinkage was calculated, technically. But when Dominic flagged that we were clearing the availability gate on three basketball candidates with thinner recent-activity records than our threshold requires, I approved the exceptions on the grounds that the corpus depth compensated. That is not how the gates work. The gates are binary for a reason.
The deeper problem was that we had stopped being suspicious of agreement. When our estimates and the lines converged, we read that as signal quality rather than as a reason to probe harder. There is a whole piece on this site about why convergence should make you suspicious of yourself, and I had apparently forgotten we wrote it.
Calling Gate Exceptions "Corpus-Adjusted Clearances"
The terminology matters here, because the terminology was part of the problem. When Dominic raised the availability concern the first time, I told him we were applying something I called a "corpus-adjusted clearance" — meaning that a player with four years of deep performance data could pass a recent-activity gate on a lighter recent record than we normally require, provided the long-run numbers were stable. I said this with enough confidence that it went into the working notes as though it were policy.
It was not policy. I had invented it on the spot to justify a decision I had already made, and then written it down, which made it look like doctrine. By week three of the month, Renata was applying corpus-adjusted clearances without checking with me, because it was in the notes. We had, in the space of about eighteen days, introduced a new gate exception, failed to stress-test it, and propagated it through the workflow — all because the results during its introduction happened to be good.
"You named it," Renata said, when we were reviewing the second grading window. "Once you name something, it starts defending itself."
She was right. The named exception acquired a kind of institutional momentum that a quiet judgment call would not have had. A quiet judgment call gets revisited. A named procedure gets applied.
The corpus-adjusted clearance was also, on reflection, a direct contradiction of the shop's oldest stated belief: sample size beats recency. We built the whole corpus methodology around the idea that a small recent sample is not a substitute for a large historical one — but the inverse is also true. A large historical sample does not substitute for sufficient recent activity. Both are required. The gates exist to enforce both simultaneously, and I had quietly suspended one of them.
What the Second Grading Window Actually Showed
The cost showed up cleanly in the calibration report. We grade every rating inside a fixed window — that is not optional and it has never been optional — and when the second window closed, the numbers were specific enough to be instructive.
The ratings that had cleared standard gates were calibrated within normal range: stated confidence matched observed accuracy to within a few percentage points, which is about what we expect in a good period. The ratings that had cleared via corpus-adjusted clearance were a different picture. The accuracy on those was not catastrophically low — this is important — but the confidence we had attached to them was substantially higher than the accuracy warranted. We had been right often enough. We had been right at a rate that did not match how confident we said we were, which is the failure mode that matters to us.
The gap was large enough that it dragged the overall calibration figure for the month into territory we would normally flag as a problem requiring a process review. We had flagged it as a problem. We were the problem.
There was a secondary cost that took longer to name. Because the good results in the first month had felt like confirmation, we had also quietly reduced the frequency of the internal disagreement sessions where Dominic or Renata push back on my gate calls. Not formally — nobody decided to do that. It just happened, because pushing back on a process that appears to be working feels contrarian rather than rigorous. This is how publishing the error rate changed how we argued — for better, mostly, but it also means that a good-looking error rate can suppress the argument you most need to have.
The honest summary: we introduced an undocumented exception, named it, propagated it, and then used the first month's results to avoid examining it. The second month examined it for us.
What the Gates Look Like Now, and What We Chose Not to Fix
The corpus-adjusted clearance is gone. It is not in the notes, it is not in the working process, and it is not available as an informal option. The recent-activity threshold is what it was before I invented an exception to it. This was not a difficult decision once the calibration report made the case plainly — the difficulty had been in seeing the case at all while the first month's results were still warm.
We also reinstated a specific practice that had lapsed: any gate exception, even a one-off, requires a written note in the grading file at the time it is made, not retrospectively. The note does not justify the exception — it just records it, so that the grading window can evaluate it explicitly rather than having it dissolve into the aggregate. Dominic suggested this years ago and we implemented it and then apparently stopped. We have started again.
What we chose not to fix is the internal disagreement cadence. I considered making it a scheduled, mandatory thing — a standing session where someone's job is to find fault with the gate calls. We discussed it. Renata was skeptical, and I think she was right to be. A mandatory disagreement session is a ritual, and rituals produce the form of the thing rather than the substance. The disagreement that caught this error was not scheduled; it came from Renata reading the grading numbers and saying something direct. Scheduling it might make it feel handled without actually handling it.
There is also something we have not resolved, which is whether the good first month was genuinely good or whether it was a run of favorable conditions that the corpus-adjusted clearance happened to survive. The accuracy was real. But we cannot now disentangle how much of it reflected the method working and how much reflected a stretch of basketball where deep-corpus players were simply performing closer to their long-run means than usual. We kept the results in the record — there is no quiet deletion here — but we annotated them. The annotation says: gate exception active; calibration weight reduced.
The thing I keep returning to is the naming. Renata's observation — that naming something lets it start defending itself — applies to more than gate exceptions. Every piece of this method has a name, and every named piece has, at some point, been applied past the boundary where it belonged. I am not sure whether the solution is fewer names or more rigorous attachment of names to explicit conditions. Probably both. Possibly neither, and the real answer is just that a good month should make you slower, not faster, and I have now written that down in enough places that I might eventually believe it.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.