PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

What Counts as Enough History

The first real decision in any evaluation is not how to score somebody. It is whether you are willing to score them at all.

We spent our first two seasons answering that badly, in the specific way that feels rigorous while you are doing it: we had a minimum, we could state it, and we had chosen it after looking at the data rather than before.

Field notes from a robot economy governed by an AI

Serialized fiction: robots, markets, and a Central Intelligence that wants everything.

Enter Mechadia

The Minimum Nobody Could Defend

Our original rule was a fixed count of prior appearances, the same across every sport. It had the great virtue of being easy to explain and the great flaw of being invented on a Tuesday.

Nobody in the shop could say where the number came from. It was not derived from anything. It sat in the config for eighteen months and every candidate in the system was measured against it, and when I finally went looking for its origin I found a commit message that said, in full, seems reasonable.

What made it dangerous was not that it was arbitrary. Arbitrary would have been survivable. It was that the same threshold behaved completely differently depending on the sport it was applied to — a fixed count is a lot of information in a sport where a player features almost every day, and almost none in one where they feature twice a week and rest a third of those.

Deriving It Instead of Choosing It

The fix we tried was to stop picking a number and start asking what the number was for.

The purpose of a minimum is to stop you making a claim your evidence cannot support. So the question becomes: at what point does adding one more prior appearance stop meaningfully changing our estimate? That is answerable, per sport, from the corpus we already had — and once we asked it that way the answer came out different in every sport, sometimes by a factor of three.

We also separated two things that had been tangled together. There is a minimum below which we will not rate somebody at all, and there is a much higher amount of history above which we start trusting the estimate rather than merely tolerating it. Those had been one number. They should never have been one number.

“You had a floor and you were treating it like a target,” Priya said, when she rebuilt the calculation. “Everything that scraped past it got the same confidence as everything with four seasons behind it.”

Two Seasons of Ratings We Could Not Stand Behind

The honest cost is that everything we produced before that fix was built on a threshold we could not justify, and a meaningful share of it was built on players who should never have been evaluated.

We did not withdraw those ratings. That was deliberate and I still think it was right — we have a rule against deleting anything, only grading it — but it means our own record contains a long stretch where the error rate is worse than it looks, because a portion of those ratings were never eligible to be made in the first place.

The subtler cost was to how we argued internally. When your threshold has no derivation, every conversation about a marginal candidate becomes a negotiation, and the person who argues hardest wins. We had months of that. I remember defending a candidate on the grounds that they were basically at the minimum, which is a sentence that should have ended the discussion in the other direction.

The Two Numbers We Now Write Down First

Every sport has a floor and a trust level, both derived, both recorded before a season starts, and both left alone once the season is running. If we want to change one we change it for next season, in writing, with the reason attached.

The part that has done the most work is the smallest: the floor is now a hard reject rather than a soft consideration. Nothing marginal gets discussed. A candidate is either eligible or it is not in the conversation, and the amount of time that has given us back is genuinely considerable.

I would not claim the current numbers are right. I would claim we can now say where they came from, which is a different and much lower bar than being right, and it is the one we failed for two years.

The commit that replaced seems reasonable is about four hundred lines longer and much less confident. That seems like the correct direction for a change of this kind.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top