PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Rules We Made Up That Turned Out to Be Wrong

Most of our rules started as a reaction to something embarrassing. We got burned by a short sample, so we wrote a rule about minimum history. We got burned by that rule being too rigid, so we amended it, and then the amendment became its own problem. The Doctrine desk exists partly to document the rules themselves, and partly to document how many of them we eventually had to retire or quietly contradict.

This piece is about four rules specifically — ones we ran on for long enough that they felt settled, ones we would have defended confidently if anyone had asked. Three of them turned out to be wrong in ways we could measure. One of them turned out to be wrong in a way we still can't fully describe, which is somehow worse.

I want to be precise about what "wrong" means here, because the shop has a specific definition. A rule is wrong when following it produces worse calibration than ignoring it would have. Not when it fails once. Not when a colleague disagrees with it. When the grading window closes and the confidence we expressed while following the rule does not match the accuracy we actually achieved. That's the test. By that test, four rules failed.

Why Everyday Things Exist

Discover the surprising reasons behind the things, rules, habits, and systems we encounter every day.

Learn more

The Rules That Felt Like Hard-Won Wisdom

The first rule was about recency caps. After a stretch where we over-weighted a player's last three performances and got badly surprised by regression, we wrote a hard cap: no single performance window shorter than twelve events could influence a rating by more than a fixed percentage. It felt responsible. It felt like exactly the kind of structural guard a careful shop should have.

The second rule was about sport-mixing. We had been running basketball and baseball candidates through the same shrinkage parameters, which was lazy, and we paid for it. So we separated them entirely — no shared priors, no cross-sport calibration, each sport treated as its own closed system. Clean. Logical. Wrong.

The third rule was about consensus. If our internal model and the line agreed within a narrow band, we treated that as a signal to reduce our stated confidence — the idea being that when everyone agrees, the interesting question is what everyone is missing. There is a version of this instinct that is correct. The rule we wrote from it was not that version.

The fourth rule was about flagging. Any candidate whose most recent event was more than a certain number of days old got a warning tag in the system, and that tag propagated downstream into the rating in ways that were supposed to represent uncertainty. It did represent uncertainty. It also, we eventually discovered, represented a silent bias we hadn't intended — one that was quietly distorting an entire category of candidates before we ever scored them.

How We Defended Each One While It Was Failing

The recency cap felt fine for a long time because it was doing real work in the cases we could see. When we reviewed the ratings that embarrassed us most, the cap had usually prevented something worse. What we weren't measuring was the category of ratings it quietly damaged — the ones where recent performance was genuinely predictive and the cap was suppressing a real signal in the name of discipline.

Remi was the one who finally ran the split. She pulled every rating from an eighteen-month window, separated them into "cap triggered" and "cap did not trigger," and compared the calibration gaps. The cap-triggered group was systematically overconfident in one direction. We had been adding a structural drag to exactly the candidates where we should have been listening more carefully.

"The rule was doing what you wrote it to do. The problem is you wrote it to solve the last mistake, and it was creating a different one the whole time."
That was Remi. She said it without any particular satisfaction, which was its own kind of rebuke.

The sport-mixing rule survived longer because the damage it did was diffuse. Separating the sports entirely meant we lost the cross-sport base rate information that had been quietly stabilizing our estimates. We didn't notice because no single rating looked obviously wrong — the error was in the aggregate, and our most accurate months were masking it. Hit rate was fine. Calibration was drifting. We were right often enough that we didn't notice we were wrong in a systematic way.

The consensus rule is the one I find hardest to explain, because the underlying instinct still seems correct to me. When agreement is total, suspicion is warranted. That's real. But the rule we operationalized from it was too mechanical — it treated narrow band agreement as a fixed trigger rather than as one input among several, and it fired in situations where the agreement was simply accurate. We were manufacturing uncertainty to satisfy a heuristic.

What Following Wrong Rules Actually Cost

The honest answer is: calibration points, and some credibility with ourselves. The shop does not act on its ratings in any market sense, so the cost was not financial. It was methodological. We published confidence levels that did not match our accuracy, and we did it while following rules we had written specifically to prevent that.

The recency cap cost us most visibly in basketball, where performance windows are shorter and recent form carries more genuine information than our rule was allowing. We had a period — about eleven weeks — where our stated confidence on basketball candidates was running roughly eight points above our observed accuracy. That's a meaningful calibration gap. We wrote about a version of this when we documented the week everything was confidently wrong, though at the time we attributed it to the wrong cause. We thought it was a corpus problem. It was partly the cap.

The flagging rule cost us something harder to quantify. When we finally traced the distortion, we found that the warning tag had been applying a confidence penalty to a whole class of candidates — specifically, players in sports with natural long gaps between events, like tennis. We had written the rule thinking about basketball, where a two-week gap is genuinely suspicious. In tennis, a two-week gap between events is routine. The rule did not know the difference. We had been under-sampling an entire sport without realizing it, and the flag was one of the mechanisms doing it.

What I keep coming back to is that each of these rules was written after a real failure. They were not invented out of comfort or laziness. They were written by people who had just been embarrassed and were trying to stop it from happening again. That did not protect them from being wrong. It may have made them harder to question, because they had a story attached — a specific miss they were designed to prevent — and that story made them feel load-bearing even when they weren't.

What Replaced Them, and What We Still Haven't Resolved

The recency cap was replaced with a sport-specific weighting function that treats the cap as a soft prior rather than a hard ceiling. It can still suppress recent form — it just has to be overridden by a larger body of evidence to do so, rather than firing automatically. Whether this is better is genuinely uncertain. It has performed better over the window we've graded it on. That window is not long enough to be confident about.

The sport-mixing ban was partially reversed. We reinstated a shared base rate for cross-sport calibration while keeping the scoring parameters separated. This was a compromise that neither Remi nor I was fully satisfied with, and we both know it might need revisiting. It is currently producing better calibration than either the fully-mixed or fully-separated approach did. We are not sure why, which is its own problem.

The consensus rule was retired without a direct replacement. We still carry the instinct — we still think total agreement deserves scrutiny — but we stopped encoding it as a mechanical trigger. It now lives in the review step as a prompt rather than in the model as a parameter. Whether that's better or just less measurable, I genuinely don't know.

The flagging rule was rewritten to be sport-aware. The gap threshold now varies by sport, calibrated against the actual distribution of event spacing in each corpus. This one I'm most confident about, not because it's clever, but because the error it fixed was concrete and the fix was concrete. It is the kind of rule that should have been written that way the first time, and wasn't, because we were thinking about one sport when we wrote it.

We kept the underlying principle behind all four rules, which is that no single performance window, no single night, and no single data point should dominate a rating without being pulled back toward the base rate. That principle survived. The specific mechanisms we used to enforce it mostly didn't.

The thing I can't stop thinking about is whether the replacement rules are wrong in ways we haven't found yet. Probably some of them are. The grading window is still open on most of them. We will find out when it closes, and we will write about it when we do — which is either a commitment to honesty or a commitment to having more material for the Doctrine desk, and I'm no longer sure those are different things.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top