What Each Sport Punishes You For
We cover five sports with substantially the same machinery, and the machinery breaks differently in each one. Not in proportion to how hard the sport is — in a specific way, characteristic of that sport, that took a season each to identify.
What follows is a list of our own weaknesses organised by which sport finds them.
Neutral explanations of government, corporate, financial, and bureaucratic systems.
We Expected Difficulty to Scale, and It Does Not
Our early assumption was that some sports would be predictable and others would not, and that our accuracy would sort itself into a ranking we could then work down.
That is not what happened. Accuracy was broadly similar across sports. What differed was the kind of error, and because we were only tracking a hit rate, the kinds were invisible. Five sports were failing at roughly the same frequency for five unrelated reasons, and the aggregate number showed a flat, uninformative line.
Basketball punished us for rest. Long, consistent records, then a player sits for reasons that have nothing to do with form, and prior performance says nothing about whether tonight is one of those nights.
Baseball punished us for the matchup. A player's own history is genuinely predictive and genuinely incomplete, because a large part of any single outing is determined by who they are facing, and we were modelling one side of a two-sided event.
Hockey punished us for role volatility. Time on ice moves sharply, week to week, for tactical reasons, and a record built on aggregate output silently mixes together two very different jobs.
Soccer punished us for rotation across competitions. The same squad plays in several, with different priorities, and a record that ignores which competition an appearance came from is averaging over decisions the coach is making deliberately.
Tennis punished us for the absence of a team. There is no rotation, no rest, no lineup — and it was still the hardest, because a single opponent determines nearly everything and there is no roster to average across.
Per-Sport Adjustments, and Refusing to Unify Them
We tried a unified correction first, naturally, because five separate adjustments felt like an admission of defeat. It failed in the way these things fail: it improved nothing and made every sport slightly worse than a sport-specific fix would have.
What we run now is one shared skeleton — corpus, gates, shrinkage, grading — and five different sets of thresholds and features hanging off it. Basketball carries an availability signal that the others do not need. Baseball carries opponent context. Hockey weights by role rather than by appearance. Soccer keys on competition. Tennis has the fewest features and the most conservative thresholds, because we trust it least.
The skeleton being shared is what makes this maintainable. The parameters being separate is what makes it work.
“You keep trying to write one thing that handles all five,” Marcus said. “The sports aren't cooperating. They're not the same problem wearing different jerseys.”
The Cost of Five Configurations
Five parameter sets mean five opportunities to overfit, and we have certainly overfit at least two of them. Tennis in particular has been tuned more times than its volume justifies, and I would not defend the current numbers as anything other than the ones that happened to survive.
It also means a change to the skeleton has to be checked five times, and we have twice shipped something that helped three sports and quietly hurt a fourth. Both times it was hockey, which has the least volume of the team sports and therefore the noisiest feedback.
And there is a maintenance cost that compounds: nobody in the shop holds all five configurations in their head. We are each better on two or three of them, which means a review is genuinely distributed and slower, and a mistake in the sport you know least well can survive a review by somebody who also does not know it well.
Tracking Error by Sport and by Kind
The change that mattered most was not a modelling change at all. It was splitting the scorecard by sport, and then within each sport by the kind of failure. A flat aggregate hid all of this for a year.
The doctrine entry is that a shared method does not imply shared parameters, and an aggregate accuracy figure across heterogeneous domains is close to meaningless — it will be stable and uninformative, which is the worst combination available in a metric.
We still refuse to unify the parameters. Somebody proposes it about once a year, usually somebody new, and the answer is a link to the season we tried it.
There is one more thing worth extracting, because it generalises past sport. Each of those five weaknesses is a case of the same underlying error: a record that averages over a decision somebody else is making deliberately. Rest, matchup, role, competition, opponent — in every case prior performance is summarising outcomes that were produced under conditions the record does not carry. The sport-specific fixes are all, structurally, the same fix: find the decision being averaged over and put it back into the data.
Tennis remains the one I would drop if we had to drop one. It has survived three reviews on the argument that the discipline of a hard sport is good for us, which may be a rationalisation.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.