Publishing the Error Rate Changed How We Argued
We publish how often we are wrong. That was originally a statement about honesty, and it is still that, but it is not what the decision has mostly done.
Its largest effect has been internal, and it was not a effect anybody predicted when we agreed to it.
Analytics, observability, and AI-driven insight from test runs.
Before, Every Argument Was About Whose Reasoning Was Better
Internal disagreements used to be settled by argument quality. Somebody proposed a change, somebody else objected, and the more articulate case won. That is a completely normal way to run a small shop and it selects for the wrong thing.
The person who argues best is not reliably the person who is right, and in a group of five the person who argues best is a fairly stable identity. Over a couple of years that produces a system shaped by one temperament rather than by evidence.
It also made changes hard to reverse. A change that had been won by argument was defended by the same argument, and reversing it required somebody to construct a better one rather than simply to observe that it had not worked. We carried at least two changes for months past the point where the results were clear.
The absence of a public number was doing a lot of work here. When nothing is committed to externally, the cost of being wrong is diffuse and social. There is always a reading of the evidence under which the change was fine.
Committing to a Figure Before the Argument Starts
What changed is not really the publishing. It is that publishing forced us to define, in advance and in writing, what result would count as the change having worked.
You cannot publish a miss rate and then privately reinterpret it, so every proposal now arrives with a prediction attached: this should move this bucket's calibration by roughly this much, within this window. That prediction is recorded before anything is deployed.
The effect on argument quality is that it barely matters any more. A proposal with a clear prediction gets tried, cheaply, and the prediction is checked. Persuasiveness has been demoted from the deciding factor to the thing that determines what gets tried first, which is a much more appropriate job for it.
“It's not that we argue less,” Dana said. “It's that the arguments have a shorter half-life. There's a date on them now.”
What It Costs to Have the Number in Public
The obvious cost is that a bad stretch is visible, and we have had bad stretches. There is a period in our record where the calibration is plainly poor for six weeks, and the explanation is that we deployed something that did not work and took longer than we should have to accept it.
The less obvious cost is that it makes us conservative in a way I am not sure is entirely healthy. A change with a wide range of possible outcomes is harder to propose when the outcome will be visible, and I suspect there are experiments we have not run because the downside would be legible. I cannot prove that — you cannot count the experiments you did not think of — but the incentive is there and it points the wrong way.
And it has not made us right more often. Our accuracy is roughly where it was. What improved was the honesty of our confidence and the speed at which we abandon things, neither of which shows up in the headline.
Predictions Recorded Before Deployment
The practice that carries the weight is the smallest one: write down what you expect the change to do, with a number and a date, before it goes anywhere. It takes a few minutes and it is the entire mechanism.
We kept the publishing too, and I would not drop it, but I have stopped describing it as the important part. The important part is having committed to something checkable in advance. Publishing is just what makes it impossible to quietly renegotiate afterwards.
Anybody could adopt the prediction habit without publishing anything. I think most people would find it uncomfortable for the same reason we did, which is that it converts a large number of confident opinions into a small number of testable ones.
The six-week stretch is still there. Somebody asked about it recently and the answer took about a sentence, which is roughly the point.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.