PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

We Graded Ourselves Generously for Months

Our error rate was too good for about five months, and the reason was not a bug and not dishonesty. It was a series of reasonable decisions that all leaned the same way.

That is the version of this failure I find most worth writing down, because there was no moment at which anybody did anything wrong.

Understanding the challenges of modern life

Calm explanations for why life feels harder than it should.

Read Why We Struggle

Four Defensible Choices Pointing One Direction

The first: a rating whose outcome was ambiguous was excluded rather than counted as a miss. Defensible — an ambiguous outcome is not evidence either way.

The second: a rating on an event that was postponed was excluded. Also defensible, and obviously correct.

The third: where an outcome arrived after our grading window, we graded it late rather than discarding it. Defensible, and it improved our data.

The fourth: near-misses on a boundary were resolved in whichever direction the source's own rounding suggested. Defensible, and it matched the convention of the sources.

Each of those was decided separately, months apart, by different people, with a sound argument. And every single one of them, on average, removed or reclassified a case that would otherwise have counted against us. Not by design. Nobody noticed the direction because nobody looked at the four together.

Counting the Exclusions Rather Than the Results

Ellen found it by asking a question none of us had: how many ratings never get counted at all, and which way do they lean?

The answer was a larger proportion than anybody guessed, and it leaned consistently. Once we counted the exclusions as a category and graded them under the least flattering reasonable interpretation, our error rate got meaningfully worse — and stayed there, which is how we know the new number is the honest one.

What we run now is two figures. One under our normal grading rules, one under the harshest defensible interpretation of every ambiguous case. The gap between them is reported. When the gap grows, that is a signal that our exclusion rules are doing too much work, regardless of what the headline says.

“Every one of these is fine,” Ellen said. “There are just four of them and they all point at you looking good.”

What made the audit possible was that we had kept the excluded cases rather than dropping them. Every exclusion was recorded with its reason, which meant the population was there to be counted once somebody thought to count it. Had we simply filtered them out at grading time — which is the obvious implementation and the one we nearly wrote — the bias would have been undetectable from the inside. There would have been no artefact to audit, only a flattering number with nothing behind it.

Five Months of a Number We Had Published

We had published the flattering figure. Not knowingly, but published is published, and the correction was public and larger than I would have liked.

The harder cost is that I cannot fully separate the innocent explanation from the convenient one. Each of those four decisions was made in good faith as far as I know. It is also true that had any of them leaned the other way, somebody would probably have questioned it, and nobody questioned these. That asymmetry in scrutiny is not neutral, and it is the closest thing to a finding about ourselves in this notebook.

I do not think we were dishonest. I think a group of people making small judgement calls about how to grade themselves will, absent a mechanism, tend to make them kindly, and the individual calls will each be defensible the whole way down.

Two Numbers, and the Gap Between Them

Every grading decision that could go either way is now recorded as such, and the harsh-interpretation figure is computed and reported alongside the normal one. The gap is the thing we actually watch.

The rule is that no new exclusion may be added without checking which direction it moves the headline. If it flatters us, it needs a stronger argument than if it does not — deliberately asymmetric, to counteract the asymmetry in scrutiny that produced this.

It also means our published error rate is worse than it would otherwise be, permanently. I would rather have the worse number and know what it means.

The gap is currently narrow. It was narrow before we started measuring it too, which is not evidence of anything and is mildly reassuring anyway.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top