PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Distrust Any Result You Are Pleased About

A bad result arrives with a question attached. A good one arrives with an answer attached, and the answer is usually that we are good at this.

That asymmetry has cost us more than any single bug, and the correction is a ten-minute habit rather than a mechanism.

Understanding the challenges of modern life

Calm explanations for why life feels harder than it should.

Read Why We Struggle

Where Our Investigation Effort Actually Went

We looked, once, at how we had spent our investigative time over about a year. Almost all of it followed bad results. That was not a policy; it was what happens by default, because a bad result is uncomfortable and discomfort allocates attention.

The consequence is that our machinery was thoroughly examined in every state that produced poor output and barely examined in any state that produced good output. Which is precisely backwards if you care about knowing whether the system works, because a broken system can produce good output — we have three documented cases of it — and every one of those went unexamined for weeks.

The most expensive was a month where a filter rejected everything and our accuracy looked outstanding. Nobody investigated. We wrote a slide.

The general shape: bad results are self-reporting and good results are self-justifying. Any process that relies on discomfort to trigger scrutiny will therefore be blind in exactly the direction that flatters it.

There is a second-order version of this that I think is worse. Because good results went unexamined, the explanations we attached to them went unexamined too — and those explanations became our account of why the method works. So our understanding of our own system was assembled disproportionately from the periods we had looked at least carefully. Every confident sentence I used to say about which components mattered came from a good month nobody had interrogated.

Ten Minutes, Asked Out Loud, Before the Discussion

The practice is one question, asked at the start of any review of a strong result: if this number is good for a reason that is not skill, what would that reason be?

It has to be asked out loud, by somebody, before the explaining starts. Once a group has begun constructing reasons the result is real, the question cannot be introduced — it reads as an attack on whoever proposed the explanation, and it will lose.

Answers that have come up: volume too low to mean anything, a gate rejecting more than it should, an unrepresentative subset, a change that has not been running long enough to have been tested by anything difficult. Three of those have turned out to be the actual cause of a good-looking result at least once.

“Ten minutes of trying to ruin it,” Ellen said, when we agreed to adopt it. “If it survives ten minutes of that, then talk about why it worked.”

It Makes Success Unpleasant

The cost is entirely to morale and it is not trivial. Nobody enjoys having their good week interrogated, and a practice that greets every improvement with an attempt to explain it away is corrosive if applied without care.

We have not fully solved that. What helps a little is that the question is directed at the result rather than at the person, and that it is asked every time rather than only when somebody is suspicious — a universal ritual is much easier to sit through than a targeted one.

It has also produced two false alarms where we spent a day chasing a non-existent artefact and the result was simply good. That is a real cost and I consider it acceptable, though I notice I am the person who has to say so and I am also the person who introduced the practice.

The Pairing With the Bad-Result Rule

The doctrine entry sits directly beside the one about investigating poor results, and the pairing is the point. Neither rule is unusual on its own. Having both, applied with equal energy, is what removes the asymmetry that was actually hurting us.

If I could keep one item from this whole notebook it would be this pairing, and not because it is clever. Because it is the only entry that addresses how we behave rather than how the system behaves, and every other failure in here was eventually traced to something about how we behaved.

The machinery can be fixed with code. This one needs somebody to ask a question in a room, every time, including on the weeks when it would be nicer not to.

One more thing about why it has to be a ritual rather than a judgement call. The whole problem is that we cannot tell from the inside which results deserve scrutiny — if we could, we would already have looked. Reserving the question for results that seem suspicious just reproduces the original bias with an extra step, because a result that flatters us does not seem suspicious. Applying it to everything is inefficient by design, and the inefficiency is the only part that makes it work.

The question has been asked at every strong-result review for over a year now. It has found something four times. I do not know what the right hit rate for a habit like that is, but four is enough that nobody has proposed dropping it.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top