The Method Transferred, the Thresholds Did Not
When we added our second sport we expected a week of configuration. The method was general, the pipeline was generic, and the new sport was structurally similar to the one we knew.
It took most of a season, and what it broke was informative in a way a smooth transfer would not have been.
The origins and reasoning behind familiar things.
Everything Ran and Nothing Was Right
The pipeline ingested the new sport without complaint. Gates fired. Scores came out. The distribution of scores looked plausible. And the grading, when it arrived, was poor in a way that had no single obvious cause.
Working backwards, nearly every number we had chosen for the first sport was wrong for the second, and each was wrong by a different amount and for a different reason. The minimum history requirement was far too low, because appearances in the new sport carried less information each. The activity window was too long, because the season had a different shape. The shrinkage was too weak, because the population variance was larger.
None of those were bugs. Every one was a parameter that had been tuned, correctly, for a different sport, and had then been silently inherited.
The uncomfortable realisation is that we could not tell, before this, which of our numbers were principled and which were fitted. They all looked the same in the config. Transferring to a second domain was the first thing that had ever separated them.
Separating the Skeleton From the Settings
What we did was mechanical and it took longer than the modelling. Every constant in the system had to be classified: is this a property of the method, or a property of a sport?
The method-level ones turned out to be few. Reject anything you cannot grade. Shrink toward the base rate. Require the event to be in the future. Grade inside a fixed window. Cap concentration. Those transferred without modification and they are, as far as I can tell, the actual content of what we do.
Everything else — every threshold, window, minimum and weight — was sport-level, and had to be derived separately for the new sport from its own corpus before anything could be trusted. That derivation is the part that took a season, because you cannot derive it without history, and building history takes as long as it takes.
“Now you know which five things you actually believe,” Ellen said, when we finished the classification. “Everything else was furniture.”
A Season of Output We Had to Discount
Every rating we produced in the new sport before the derivation was built on inherited settings, and we graded all of them, and they are in the record. That stretch drags our overall calibration down and it should.
We also learned something unflattering about how we had been describing our own work. I had, in good faith, characterised the method as sport-agnostic. It was not — it was one sport's parameters wearing a generic wrapper, and the wrapper was the part I had been pointing at.
The other cost is that we now know a third sport will take a season too. That is priced in and it has made us much slower to add coverage. I think that is correct and I also notice it functions as a reason not to try things, which is the same shape as an excuse.
Every Constant Is Labelled by Scope
Every constant in the configuration now carries its scope — method-level or sport-level — and a sport-level constant cannot be read without specifying which sport. That makes inheritance impossible rather than merely discouraged, which is the only version of this rule that survives contact with somebody in a hurry.
The five method-level rules are written down separately, in one short document, and they have not changed since. When somebody proposes a change to the method, the first question is whether they mean the method or one sport's settings, and about half the time asking resolves it.
I would recommend the exercise to anybody with a system that has only ever run in one domain. You will find out how much of what you believe is a belief and how much is a fitted number, and the ratio is likely to be worse than you expect.
The classification exercise had a side effect I would not have predicted. Once the method-level rules were listed on their own, it became obvious that none of them are about sport. They are about epistemics — what you are allowed to claim, how much you should believe a thin sample, when a claim becomes checkable. Which means the transferable part of what we do is not sports analysis at all. It is a small set of rules about not overstating evidence, and the sport is where we happen to practise them.
The one-page document is the most valuable thing we produced that year. It is shorter than this article.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.