PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Write the Threshold Down Before You See the Data

Choose the number first. Then look. Not the other way round, not iteratively, not "choose a rough one and refine it against the results".

It is the oldest rule in our doctrine and the one people argue with most, including me, for about a year.

Understanding the challenges of modern life

Calm explanations for why life feels harder than it should.

Read Why We Struggle

Every Threshold We Had Was Chosen After the Fact

When we audited our configuration we found that essentially every number in it had been arrived at by trying values and keeping whichever produced the best results.

That sounds like tuning, which is a respectable activity. The problem is what it does to your ability to make a claim. A threshold selected because it produced good outcomes on the data you have is not a threshold — it is a summary of that data. Applying it to new data and reporting the result as though the threshold were independent is a straightforward error, and we had made it in about a dozen places simultaneously.

The symptom, which took us a long time to connect to the cause, was that our performance was reliably a little worse than our testing suggested. Not dramatically. Consistently. Every time we deployed something it underperformed its evaluation by a similar small margin, in every sport, for two years.

That margin was the cost of having fitted our thresholds to the same history we then evaluated against. We had a persistent, measurable, unexplained gap sitting in plain sight, and the explanation was in how we picked our numbers.

Committing In Writing, With a Date

The practice is unglamorous. Before touching data, write down the threshold, the reasoning, and what you expect it to do. Date it. Then run it, once, and record what happened.

If it performs badly you do not adjust and re-run. You record that it performed badly and, if you want to try a different value, that is a new commitment with a new date. The prohibition on re-running against the same data is the entire mechanism — without it you are back to fitting, just more slowly and with better documentation.

What this bought us immediately was that the gap closed. Our deployed performance now lands about where the evaluation said it would, which is the first time in the shop's existence that has been true.

“You're not allowed to be less wrong,” Marcus said, when we were arguing about it. “You're allowed to know how wrong you're going to be. Those are different products.”

It Makes You Look Worse and Feel Slower

Our evaluated numbers got worse the moment we adopted this, because they stopped being fitted. Anybody comparing our reported figures across that boundary would conclude the method deteriorated. It did not; the measurement became honest.

It is also slower in a way that is genuinely frustrating. You get one shot per commitment, and if you chose badly you wait for more data rather than iterating. There have been two occasions when I was fairly sure a threshold was wrong within a week and had to leave it in place, producing worse output, until the window closed.

And I broke the rule once. A threshold I had committed to was plainly poor, and I changed it mid-window and re-ran, and the result was good, and I did not report the re-run for about a fortnight. That is the least defensible thing in this notebook. It is here because leaving it out would make the doctrine look easier to follow than it is.

The Commitment Log

There is a file. It has a date, a number, a reason and an expectation on each line, and a result appended when the window closes. It is the least sophisticated artefact we maintain and the one I would rescue first.

What it mostly does is make the fitting impossible to do accidentally. Nobody sets out to fit a threshold to their evaluation data; you do it by trying something, seeing a better number, and keeping it, which does not feel like a methodological choice at any point along the way. Writing the commitment down first is what converts it into a choice you would have to make deliberately.

The log also has the two occasions where I wanted to change something and could not, with what happened, and the one occasion where I did it anyway.

The gap between evaluated and deployed performance is currently about as small as it has ever been. I attribute nearly all of that to a text file.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top