PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Our Most Accurate Month Was Our Worst Result

The best month in our record is a month I would like to remove from it, and cannot, because we do not remove things.

Our accuracy that month was the highest it has ever been by a clear margin. The reason is that the system was broken in a way that happened to be flattering.

Where test results become engineering insights

Analytics, observability, and AI-driven insight from test runs.

Read iTestResults

A Very Good Number With No Explanation Behind It

The first sign was volume. We produced far fewer ratings than usual, which we attributed at the time to a thin schedule. Then the grading came in and the hit rate was remarkable.

Nobody was suspicious for about ten days. A good result arrives with its own justification attached — we had recently changed two things, and the obvious reading was that one of them had worked. We spent a week and a half discussing which.

What was actually happening is that an availability check had stopped matching, so nearly everything was being rejected. The handful of ratings that survived came through a narrower path with different characteristics, and that subset was both small and unrepresentative. A small unrepresentative sample of anything will sometimes produce a spectacular number.

So the excellent month was excellent for the same reason a coin can land the same way five times. And we had spent ten days building explanations for it, all of them plausible, none of them connected to the cause.

Making Volume Part of Every Result We Report

The mechanical change was small: no accuracy figure is reported without the count it was computed from, and no monthly figure is discussed at all below a minimum count. Below that threshold the honest statement is that we do not have a month, not that we had a good one.

The harder change was procedural. We now investigate unusually good results with the same energy as unusually bad ones. That sounds obvious. It is not what anybody does by default, ourselves very much included — a bad result demands explanation and a good one supplies its own.

What that looks like in practice is a short standing question: if this number is good for a reason that is not skill, what would that reason be? Ten minutes, at the start of any review of a strong result. It has caught two further things, both smaller than the availability failure and both real.

“We'd have shipped that month as a win,” Tobias said afterwards. “We had the slide written.”

There is a structural point here that outlasts the specific incident. A hit rate is a ratio, and a ratio computed from a small numerator is not a weak measurement — it is a measurement of something else entirely. We had been treating volume as a matter of productivity, something to report alongside quality, when it is actually a precondition for quality being measurable at all. Below a certain count there is no such thing as a good month or a bad one; there is only a month you cannot say anything about, and reporting a figure for it creates an impression that no amount of surrounding caveat undoes.

The Ten Days, and What They Say About Us

The direct cost was ten days of attention spent explaining a result that had nothing to explain, and a month in our public record that looks better than our method was.

The indirect cost is the one that has stayed with me. During those ten days we were reasoning carefully, using real evidence, and arriving confidently at conclusions that were entirely wrong. Nobody was sloppy. The care was genuine and it was pointed at a fiction. That is a much more frightening failure mode than carelessness, because there is no personal quality you can cultivate to prevent it — only a procedure.

We also nearly kept one of the two changes on the strength of that month, and it was the worse of the two. It survives in the notebook as the closest we have come to institutionalising a coincidence.

Symmetric Suspicion

The doctrine entry is one line: distrust any result you are pleased about. It sits directly next to the entry telling us to investigate bad results, and the pairing is deliberate — the asymmetry between how we treated the two was the actual defect.

The month stays on the record with a note attached explaining what the system was doing. That note is the most-read thing we have ever written internally, mostly by us, mostly when somebody is about to get excited.

I would not claim we are now immune. I would claim we have a ten-minute habit that has caught three things, which is three more than the previous approach of feeling good caught.

Our best month remains, on paper, our best month. Everybody here knows what it was and nobody cites it, which is about the best outcome available once the number exists.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top