The Number That Turned Out to Measure Nothing
There was a column in our scoring that everybody trusted, nobody had built, and which turned out to be a very slightly noisy copy of another column.
It survived eleven months. It survived because it behaved exactly as a good feature behaves: it correlated with outcomes, it moved when you expected it to move, and removing it made the model worse.
Neutral explanations of government, corporate, financial, and bureaucratic systems.
A Feature That Passed Every Test We Knew How to Run
It was a composite — a blend of a hit rate and a consistency measure, computed over a rolling window. It had been added early, by somebody reasoning sensibly, and it did well in every check we applied.
It correlated with graded outcomes. Removing it degraded performance. Its distribution looked healthy. It varied across players in a way that matched intuition. On any dashboard you would have called it one of our better inputs.
What none of those checks could see is that it was almost perfectly determined by a simpler feature we already had. It was not adding information; it was restating existing information in a slightly noisier form. Removing it degraded performance because the model had learned to lean on it, and taking away half of a redundant pair briefly hurts before anything readjusts.
That is the trap in a sentence. Redundancy looks like predictive power from the outside. Both features move with outcomes, so both look useful, and correlation with the target is exactly the wrong test for asking whether a feature earns its place.
Asking What Each Input Adds Rather Than What It Predicts
The check we were missing is embarrassingly standard: for each input, how much does it improve things given everything else we already have. Not in isolation. Conditionally.
Run that way, the composite added essentially nothing. Its conditional contribution was indistinguishable from noise, and the simpler feature it duplicated accounted for effectively all of the apparent value. We removed it, allowed the model to resettle, and performance returned to where it had been and then very slightly past it.
The slight improvement is the part I want to note, because it is the part I would not have predicted. Carrying a redundant noisy input is not free even when it looks harmless — it gave the scoring an extra way to be confident, and some of that confidence was manufactured from the noise rather than the signal.
“It was answering a question you already had an answer to,” Ellen said, “and answering it slightly worse.”
Eleven Months of Confidence We Had Not Earned
The ratings from that period were not badly wrong. They were, on average, slightly more confident than the evidence supported, because a redundant input was inflating agreement between components that were not actually independent.
That shows up in the calibration record rather than the accuracy record, which is precisely why it took so long to notice. We were about as accurate as we should have been. We were consistently overconfident by a small margin, and a small consistent margin is nearly invisible month to month and unmistakable across a year.
The other cost is that I had cited this feature, more than once, when explaining to people why our scoring worked. I described a mechanism that was not doing anything. Nobody was harmed by that beyond my own credibility, but it is a good illustration of how comfortably a plausible explanation attaches itself to a component that is merely present.
Every Input Now Has to Justify Itself Against the Others
Nothing enters the scoring without a conditional contribution check, and every existing input is re-checked at the end of each season. Two more have been removed since. Neither removal hurt.
We also stopped treating "removing it makes things worse" as evidence that a feature is valuable. That test cannot distinguish between a useful input and a redundant one the model has come to depend on, and it was the only test we had for the better part of a year.
The broader habit is to be suspicious of any component nobody can trace the origin of. This one had no author anybody remembered and no written justification, and it turned out to be the one doing nothing — which is weak evidence but I now weight it more than I used to.
The simpler feature it duplicated is still in there, unchanged, doing the whole job on its own. It has a comment above it now explaining what happened to its more elaborate sibling.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.