PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

Individual Sports Break Our Method

The first time we tried to apply our corpus-and-gates process to a tennis player, we produced a rating in about four hours and felt very good about it. The rating was wrong in the specific way that confident estimates are wrong: it was precise, internally consistent, and measuring something that didn't matter. We had built a body of evidence about what the player had done across dozens of matches, run it through every gate, shrunk the estimate toward the base rate, and graded it against outcomes. The architecture was sound. What we had missed was that in an individual sport, the architecture is only half the problem.

In team sports, noise is partially absorbed by the structure around the player. A basketball player who has an off night still exists inside a system — a rotation, a set of teammates, a scheme — that distributes performance across five people and dozens of possessions. The corpus can hold a lot of that signal because the player's role is relatively stable. The environment cooperates with the measurement. Individual sports offer no such courtesy. The player is the whole unit. Every variable that a team would diffuse lands entirely on one person, and the corpus has to carry it alone.

We spent the better part of one season learning this the hard way, and what follows is an account of that, not a prescription for fixing it.

The stories behind the things around us

The origins and reasoning behind familiar things.

Read Why This Exists

The Corpus Assumption That Team Sports Had Been Hiding

When we built our original corpus framework, we built it on a quiet assumption we had never articulated: that a player's historical output was roughly independent of their opponent's identity. Not perfectly independent — we modeled matchup adjustments, strength-of-schedule corrections, all of it. But the assumption underneath those corrections was that the adjustments were small, that the player's base rate was stable enough to be the dominant term in the estimate.

In basketball, that assumption holds well enough. A player who averages a certain number of points per game against the league will be within a reasonable band against any specific opponent, and the adjustments are real but not catastrophic. The base rate earns its authority.

In tennis, the assumption collapses. The entire sport is a direct confrontation between two people whose styles interact in ways that can swing outcomes far outside any base rate. A player who is dominant against flat hitters can be structurally exposed by heavy topspin. A player with an exceptional second serve faces a completely different set of pressures against a returner who specializes in that shot. These are not small adjustments. They are, in some matchups, the whole story. Our corpus was measuring the player in aggregate, but individual sports are played in specific, and the gap between those two things was wider than we had accounted for.

Riya, who handles most of our tennis work, said something about this that I wrote down at the time.

"In basketball you're measuring a player against a distribution. In tennis you're measuring a player against a person. Those are not the same problem and the corpus doesn't know the difference."

She was right, and we had been acting as though she wasn't.

What We Built to Close the Matchup Gap

The obvious response was to build matchup-specific sub-corpora. Instead of a single body of prior performance per player, we would maintain a layered structure: aggregate history, surface-stratified history, and then opponent-style history, where style was defined by a handful of measurable characteristics — serve speed, rally length tendency, net approach rate — rather than by the opponent's identity directly. Identity changes; style is stickier.

This felt like the right architecture. It addressed the thing Riya had identified. It also required us to have enough data in each sub-corpus to say anything meaningful, which turned out to be the second problem hiding behind the first.

For established players with long records, the sub-corpora were viable. We could split a four-season corpus into surface and style buckets and still have enough observations in each bucket to clear our minimum sample threshold. For younger players, or players who had recently changed their game, the buckets were thin. A player who had only played eight matches on a particular surface against a particular style of opponent gave us eight data points to work with, and eight data points is not a corpus. It is a rumor.

We ran the layered approach for one full season across tennis and also attempted a version of it for a hockey player whose performance was heavily dependent on line combinations — an individual-within-a-team problem that shares some structure with the matchup issue. The hockey version worked better, partly because line combinations are more stable than opponent styles across a season, and partly because hockey's team structure still provided some of the diffusion that individual sports lack entirely.

Where We Were Wrong, Specifically

The layered corpus produced better ratings for the players where the sub-corpora were thick. For the players where the sub-corpora were thin, it produced worse ratings than the original flat approach, and we did not catch this until the end-of-season calibration review.

The mechanism was straightforward in hindsight. A thin sub-corpus is more susceptible to recency bias than a flat corpus, because there are fewer observations to anchor the estimate. When we stratified by style and a player's last three matches against that style had gone badly, the sub-corpus moved toward that recent evidence much faster than the flat corpus would have. We had built a system that was more sensitive to the thing we cared about — matchup specificity — but also more sensitive to noise. The two sensitivities came bundled together, and we had not separated them.

Our calibration check showed it plainly. For players with sub-corpus sizes above our standard gate threshold, stated confidence matched observed accuracy reasonably well. For players below that threshold, we were overconfident by a margin that was embarrassing to write down in the review document, and I will not reproduce the number here because it would not tell you anything useful about the method. What it told us was that we had applied a sophisticated structure to a data problem without first confirming that we had enough data to support the structure. That is a beginner's error, and we made it.

Tomás, who runs our calibration reviews, did not say "I told you so," but he had flagged the thin-bucket risk before the season started and I had noted it and moved on. That is on me.

What the Process Looks Like Now, After the Revision

We kept the layered corpus structure, but we added a gate that did not exist before: a sub-corpus size check that runs before the matchup adjustment is applied at all. If the relevant style bucket does not clear a minimum observation count, the rating falls back to the flat aggregate and the matchup adjustment is flagged as unavailable rather than estimated from thin air. The rating is less specific in those cases. We decided that less specific and honest was preferable to specific and wrong.

The gate has a cost. It disqualifies a meaningful number of individual-sport candidates who would have passed under the old system. Early-career players, players returning from long absences, players who have recently changed surface preference — all of them hit the gate more often than they used to. The shop evaluates fewer individual-sport players than it did two years ago, and the ones it does evaluate have longer, more stable records. That is a real narrowing of scope.

What we did not change was the fundamental shrinkage principle. If anything, we tightened it for individual sports. Because the matchup interaction can produce genuine outlier performances that are structurally explainable but statistically extreme, we pull individual-sport estimates harder toward the base rate than we do for team sports. An outlier in a team sport is often noise. An outlier in an individual sport is sometimes real — but we have found that assuming it is noise and being occasionally wrong in that direction is less costly than assuming it is signal and being wrong in the other direction. The calibration history supports that, at least so far.

The broader lesson — if it is a lesson, which I am not certain it is — is that individual sports surface assumptions that team sports had been quietly validating for us. The corpus works in team sports partly because team sports are built in a way that makes corpora work. That is not a property of the method. It is a property of the sport.

We still do not know whether the matchup-style framework, given enough seasons of data, will close the calibration gap fully or whether individual sports have a structural ceiling on what any corpus-based approach can do. Riya thinks the ceiling is real and we are already near it. Tomás thinks we are not measuring the right style dimensions yet. I am not sure either of them is wrong.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top