PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

When Everyone Agrees, Be Suspicious of Yourself

There is a version of this problem that feels like confirmation and a different version that is confirmation, and for a long stretch last winter we could not tell them apart. We were looking at a basketball player — invented here, as always, to keep the mechanism visible — whose performance corpus, whose recent activity, and whose implied expectation from the published line all pointed in the same direction. Three sources. Clean agreement. We moved quickly, which is the first sign something has gone wrong.

The shop's method is built around the idea that a line is evidence about someone else's beliefs, not an instruction. When you read it that way, agreement between the line and your own model is interesting: it either means you and the market have independently arrived at the same place, or it means you have been reading the same inputs and dressed them up as independent thought. The difference matters enormously and is almost impossible to feel from the inside. Agreement is comfortable. It does not announce itself as a warning.

We did not catch it in time. What looked like convergence was a loop: our corpus features were downstream of the same recent-game data that had moved the line, and the recency signal we thought we were adding separately was just the same number wearing a different coat. Three sources became one source. The rating survived the gates, cleared calibration thresholds, and was wrong in a way that was entirely our fault.

The house doesn't gamble

Victor Draemont on judgment, patience and structural advantage — from the casino floor up.

Read Haus Edge Capital

The Specific Geometry of Unanimous Agreement

The problem is not that agreement is bad. Sometimes the corpus, the recent signal, and the line genuinely do converge from different angles, and that convergence is informative. The problem is that the three inputs have to be actually independent for the agreement to mean anything, and independence is harder to verify than it sounds.

In this case, the corpus feature we were using was a rolling performance metric weighted toward the last eighteen games. The recency signal was a five-game form indicator. The line, we later established, had moved in response to the same five-game run. So what we had was: a long-window metric being pulled by a short-window run, a short-window indicator measuring that same run directly, and a market that had already repriced in response to the run. Presented as three columns. Actually one column, justified three times.

Marta, who maintains the corpus infrastructure, put it plainly when we were reconstructing the failure afterward:

"You didn't have three signals agreeing. You had one event propagating through three different containers. The containers looked different so you counted them separately. That's a feature correlation problem, not a model problem — but it produces the same result."

She was right. The containers — the rolling metric, the form indicator, the line — each had their own label and their own section of the input sheet. Nothing in our process at that point required us to ask whether their variance was shared. We had gates for availability, for sample size, for recency adequacy. We did not have a gate that asked: where did each of these signals actually learn what it knows?

Adding a Provenance Check to the Gate Sequence

The fix we tried was to add what we started calling a provenance check — a step, sitting just before the final score is assembled, that asks each input feature to name its primary data source. If two or more features share a source, they are flagged as potentially correlated and the confidence on the combined rating is deflated before it reaches calibration.

In practice, this meant building a small dependency map for every feature in the corpus. Rolling metrics note which game windows they draw from. Form indicators note their lookback period. Line-derived features note the date and context of the movement we are reading. When the windows overlap substantially — we used a threshold of sixty percent shared games — the system marks the pair and reduces the effective input count.

The deflation is not dramatic. We are not zeroing out correlated features; correlated features still carry information. The adjustment is more like the shrinkage step that already exists elsewhere in the method: pull the confident number back toward the base rate, because confidence assembled from redundant parts is not the same as confidence assembled from independent ones. The rating can still clear the threshold. It just has to do so with a more honest account of what is actually behind it.

We ran the provenance check backward through three months of prior ratings to see what it would have changed. It flagged roughly one in five as carrying inflated confidence from overlapping sources. Most of those ratings had been directionally acceptable — the overlap had not always produced a wrong answer. But several of the worst misses from that period were concentrated in the flagged group, which was at least consistent with the hypothesis.

What the Provenance Check Got Wrong Immediately

The first version of the provenance check was too aggressive, and we knew it within two weeks. The sixty-percent overlap threshold was chosen without much principled reasoning behind it — it felt conservative, which is not the same thing as being correct. What it actually did was flag a large number of ratings where the shared source was the corpus baseline itself: the multi-season performance history that underlies almost every feature we build.

Of course most features share the corpus baseline. That is the point of having one. Flagging that shared lineage as a correlation problem was like penalizing two different measurements of the same river for both containing water. The provenance check, in its first form, could not distinguish between features that are correlated because they are measuring the same recent event and features that are correlated because they are both grounded in a long shared history. Those are not the same problem.

We spent the better part of a month rebuilding the dependency map to separate baseline lineage from event-window overlap. The revised version only flags features whose variance — not just their source — is dominated by the same short window. A rolling metric that moves primarily because of a five-game run and a form indicator that moves for the same reason are flagged. A rolling metric that is stable across the last sixty games and a form indicator that happens to sit inside that window are not.

The cost was time and, briefly, a period where we had a provenance check running that was making things worse rather than better. We kept grading through it, which is the only way to know. The error rate in the flagged group improved after the revision. It did not disappear.

What the Rating Process Looks Like Now, at That Step

The provenance check is now the fourth gate, sitting between the recency-adequacy check and the final scoring assembly. It runs automatically and produces a flag — not a rejection, a flag — with a note describing which features share variance-dominant windows and by how much. The scorer sees the flag before the rating is finalized.

What changed more durably, though, is a habit. Before any rating where the corpus signal, the recent signal, and the line are all pointing the same direction, we now ask the question explicitly: did these arrive here from different places, or did they all learn from the same recent games? Sometimes the answer is that they genuinely are independent, and the agreement is real. Sometimes the answer is that we built an echo chamber with three rooms and called it a consensus.

The deeper issue, which the provenance check does not fully solve, is that unanimous agreement is the state in which we are least likely to ask skeptical questions. The method has shrinkage built in for confident estimates. It has caps built in for concentration risk. What it did not have — and still only partially has — is a mechanism for recognizing when confidence is high specifically because the inputs are agreeing, and then treating that as a reason for suspicion rather than reassurance.

We also kept something that predates the whole episode: the rule that sample size beats recency. The failure described here was partly a failure to honor that rule. The five-game run had moved the line, moved our form indicator, and pulled the rolling metric. We let recency propagate through the system because it was dressed up as consensus. The corpus, properly weighted, would have been more skeptical. We just were not listening to it carefully enough at the time.

There is probably a version of this problem that is impossible to fully engineer away — the one where you have genuinely independent signals that all happen to be wrong in the same direction, for reasons the corpus has not seen before. The provenance check addresses the case where agreement is illusory. It does not address the case where agreement is real and the world is about to do something new. We are not sure those are separable from the inside, either.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top