A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

When Two Sources Agreed and Both Were Wrong

The working assumption, for most of the time we have been doing this, was that disagreement between sources was the thing to worry about. If two published expectations for the same player on the same night landed far apart, that was the signal to slow down, check the corpus, and ask what each source knew that the other did not. We wrote a whole piece about what disagreement between sources tells you, and we believed it. Disagreement meant uncertainty. Agreement meant the market had settled.

That belief held up well enough that we stopped examining it. Which is, in retrospect, exactly when it became dangerous.

Late in a hockey stretch last season, we were rating a forward whose per-game production numbers had been steady across a long corpus — steady enough that we had real confidence in the estimate. Two separate sources published nearly identical expectations for his next outing. They were within a fraction of a unit of each other. We treated that convergence as independent confirmation. We filed the rating, logged the confidence, and moved on. What happened next was not a small miss. It was the kind of miss that makes you go back and reread your own assumptions from the beginning.

Discover How the Systems Around You Really Work

Understand the government, financial, healthcare, business, and technology systems affecting everyday life.

Learn more

The Specific Thing Convergence Was Hiding in This Case

The two sources agreed because they were drawing from the same upstream data. Not the same vendor, not the same interface — but the same underlying game-log feed, which had a quiet versioning problem we did not know about. One source had ingested it directly. The other had aggregated from a third party that used the same feed. They looked independent. They were not.

This is a variant of the problem we ran into when the corpus we built around the wrong version of a stat quietly corrupted several months of ratings before anyone noticed. In that case the error was in our own data. Here, the error was upstream of two sources we had assumed were checking each other. The convergence was not two people independently arriving at the same conclusion. It was one conclusion, laundered through two pipelines until it looked like two.

The player in question had missed two practices that week with a minor soft-tissue issue. That information existed — it had been reported in a short item on a regional site we did not monitor. Neither source had it. Neither source's expectation reflected it. They agreed, precisely and confidently, on a number that described a player who was fully healthy and fully featured, which he was not.

What We Did When the Rating Came Back Wrong

The first thing we did was the thing you always do: we checked whether the corpus was at fault. The forward's long-run numbers were solid — over ninety games, which is the kind of sample we trust. We ran the gate checks again. We looked for a stat-type contamination of the kind that had burned us before. Nothing surfaced. The corpus looked clean. The gates looked clean. The rating looked, by our own process, correct.

It took Renata about forty minutes of source archaeology to find the feed overlap. She pulled the ingestion timestamps from both sources, traced the game-log version each had used, and found they resolved to the same build. "They're not two sources," she said, dropping a screenshot into the shared folder. "They're one source with two front doors."

"The agreement wasn't evidence. It was an echo. We were listening to ourselves and calling it confirmation."
— Renata

Once we understood the mechanism, we went back through the previous eight weeks looking for other instances where we had logged high confidence partly on the basis of cross-source agreement. We found six ratings where the two sources we had cited shared the same upstream feed. Three of those had graded well. Three had not. The hit rate on that subset was not better than chance, which is what you would expect if the confirmation was illusory.

What the Confidence Log Actually Showed We Got Wrong

The calibration damage was specific and uncomfortable. In the window covering those six ratings, our stated confidence was higher than our observed accuracy by a margin that, in any other context, we would have flagged immediately. We had been calling these convergence-supported ratings our most reliable. The calibration record said they were average at best.

The honest version of what happened is that we had a rule — cross-source agreement raises confidence — and we had never stress-tested whether the sources were actually independent. Rules we made up that turned out to be wrong is a category we have written about before, but we tend to catch those rules when they fail loudly. This one failed quietly, across a range of ratings, and only became visible when we went looking for it after a single bad grade.

There was also a secondary error, which was more embarrassing. When Renata first flagged the feed overlap, I assumed it was a narrow, one-time problem. I was wrong about that too. The third-party aggregator that had laundered the feed had been in our source list for nearly a full season. Every rating that cited both that aggregator and the direct feed had the same structural problem. We had been treating a single data source as two for months without knowing it.

Dario, who manages the source registry, put it more plainly than I would have: "We were counting the same witness twice and calling it corroboration."

The Check We Added and the Rule We Retired

We retired the confidence boost for cross-source agreement. That was not a difficult decision once we understood what had happened. Agreement between sources now triggers a provenance check before it counts for anything — we trace each source to its upstream feed and confirm they are genuinely independent before treating their convergence as evidence. If we cannot confirm independence, we treat the agreement as a single data point, not two.

The provenance check is slower. It adds a step to a process that was already not fast. We have missed the occasional rating window because the check took longer than the available time. That cost is real, and we have not found a way to eliminate it. What we have found is that the check surfaces feed overlaps more often than we expected — roughly once every three weeks since we implemented it, which is a higher frequency than I would have predicted before the hockey miss.

We also tightened the availability gate. The forward's practice status had been reported and we had not seen it, which is a gate failure rather than a corpus failure. We added a second monitor for the regional sites that cover individual team schedules, specifically because those are the places where soft availability information tends to appear before it reaches the larger aggregators. This is unglamorous work. It is also, in our experience, where the real gate leaks are — not in the model, but in the information that never reaches it. This echoes something we learned during the week everything was confidently wrong, when a cluster of bad grades traced back not to bad modeling but to availability information arriving after we had already filed.

The corpus for this player was not the problem, and we did not change it. Ninety games of genuine history is not what failed here. What failed was the belief that two sources saying the same thing meant two people had looked at the same player and reached the same conclusion independently. Sometimes it means that. Sometimes it means one person looked and the other one copied.

The unsettling part is not that the sources shared a feed. It is that we had no procedure for checking, because we had never seriously imagined they would. Agreement felt like safety. I am not sure what agreement feels like now, except that it feels like a question rather than an answer.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top