The Corpus Entry We Kept Updating
At some point last winter, Dara noticed that one of our corpus entries had been touched eleven times in six weeks. Not corrected — touched. Small adjustments to the trailing window, a reweighted season, a note appended and then removed. The underlying player profile had not changed in any meaningful way. The performance record was the same. What had changed, repeatedly, was our interpretation of it, and we had not once written down why.
The entry was for a basketball player — invented for this piece, as all our examples are — whom I'll call the Forward. He had seven seasons of reliable corpus data, a stable statistical profile, and a recent four-game stretch that looked nothing like the previous seven years. We had not changed the rating. But we kept opening the file. We kept adjusting the weights at the margins. We kept, in other words, trying to make the corpus say something it wasn't saying, without admitting that was what we were doing.
This is the piece about that. It is not a flattering piece.
Create short links, track clicks, and understand your audience. Privacy-friendly by design. No cookies, no tracking pixels, just the stats you need.
Why We Kept Opening a File That Didn't Need Opening
The Forward's four-game stretch was genuinely unusual. His output in the category we track most closely had dropped by roughly a third. Not gradually — sharply, across consecutive performances, with no obvious structural explanation. No reported availability issue. No change in role that we could confirm from the sources we use. Just a cliff, and then silence from the record about why.
The corpus, properly read, said: this is noise. Seven seasons of data does not get overwritten by four games. That is not a principle we invented — it is the oldest thing the shop believes, and the one we have tested most. We have been in situations before where the corpus was right and something else in our process was wrong, and the corpus was right then too. Sample size beats recency. We knew this.
And yet. The file kept getting opened. Dara opened it. I opened it twice. Tomás, who handles the basketball corpus almost exclusively, opened it and changed a single decimal in a trailing average before changing it back. None of us logged a reason. The entry sat there accumulating edits that cancelled each other out, like a document that everyone is afraid to close.
What We Told Ourselves We Were Doing
The official story, if you had asked any of us at the time, was that we were monitoring an anomaly. That is a legitimate thing to do. Anomalies are worth watching. A four-game drop with no structural explanation is exactly the kind of signal that sometimes precedes a gate disqualification — an availability issue that hasn't been reported yet, a role change that shows up in the numbers before it shows up in the news. Most candidates that eventually die at a gate show early warning signs in the corpus, and we take that seriously.
So the monitoring story was not entirely wrong. It was the frequency that gave us away. Eleven edits in six weeks is not monitoring. Eleven edits is anxiety dressed up as diligence. We were not watching the entry for new information. We were returning to it hoping the new information would arrive and resolve the discomfort of not knowing.
Tomás, when Dara finally put the edit log on the table, said something that I wrote down immediately:
"We weren't updating the corpus. We were visiting it. There's a difference, and we kept pretending there wasn't."
That is a precise description of what happened. Visiting a corpus entry is not a method. It is a feeling masquerading as one.
What the Repeated Edits Actually Cost Us
The direct cost was small and embarrassing. Two of the eleven edits introduced minor inconsistencies in how we had weighted the trailing window — not enough to change the rating, but enough that if we had published anything referencing the Forward's entry during that period, the internal record would not have matched the published claim. We would have been unable to grade it cleanly against what we had actually believed at the time of publication. That is a failure mode the shop cares about more than most.
The indirect cost was harder to quantify but more interesting. We had, without deciding to, created a two-tier corpus: entries we trusted and entries we kept visiting. The Forward's entry had been quietly demoted to the second tier not because the data warranted it, but because the recent stretch made us uncomfortable. The seven seasons hadn't changed. Our relationship to them had.
What we got wrong, specifically, was conflating the discomfort of unexplained variance with evidence that the corpus needed revision. They are not the same thing. Unexplained variance is normal. It is everywhere in the record. The corpus is not a document that explains things — it is a document that counts them. We had started asking it to do something it was never designed to do, and when it couldn't, we kept editing it as though more edits would help.
There is a version of this error that we have written about before in a different context — the corpus being wrong not because the data is bad but because we kept intervening in it. The mechanism is the same. The intervention feels like maintenance. It is not.
What We Changed in How the Corpus Gets Closed
We added a field. It sounds minor. It has been the most consequential single change we made to corpus maintenance in the past year.
Every entry now requires a logged reason for any edit made after the initial build. Not a long reason — a sentence is enough. But it has to be written down at the time of the edit, not reconstructed afterward. "Correcting a data entry error in game 14 of season 4" is a valid reason. "Checking on the recent stretch" is not an edit reason; it is a visit, and visits do not belong in the edit log.
The rule sounds obvious. We did not have it before the Forward's entry surfaced the problem. Dara proposed it the same afternoon Tomás named the distinction between updating and visiting, and we implemented it within the week. In the three months since, the Forward's entry has been opened twice. Both times for legitimate corrections. Both times with a logged reason.
We also added a standing review of any entry that receives more than three edits in a rolling thirty-day window. Not to investigate the player — to investigate ourselves. The question the review asks is not "what is happening with this entry" but "why do we keep coming back to this entry." Those are different questions, and only the second one is honest about what the edit log is actually measuring.
The Forward's rating, for what it is worth, has not changed. The seven seasons still say what they said. The four-game stretch resolved into noise, which is what the corpus said it was, which is what we should have trusted in the first place.
I am not sure whether the lesson here is about recency bias, or about the difference between discomfort and evidence, or about something more basic — that a corpus is a record and not a conversation, and we kept trying to have a conversation with it. Probably all three, and probably none of them fully captures it. What I keep coming back to is Tomás's distinction: visiting versus updating. I wonder how many of our other entries are being visited right now, by someone who would describe it as monitoring, and whether the edit log would know the difference.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.