Rotation, Rest, and the Limits of Prior Performance
The corpus is built on what a player has done. The problem, which took us longer than it should have to articulate clearly, is that "what a player has done" is not a fixed quantity. It shifts depending on how often the sport asks them to do it, how much recovery sits between each asking, and how much the coaching staff treats their minutes as a managed resource rather than a given. Those three variables do not behave the same way across sports. When we tried to apply the same prior-performance logic we had refined in basketball to hockey, and then to soccer, the method transferred — but not cleanly, and the seams showed up in the grades.
The specific episode that prompted this piece: a hockey forward whose corpus was deep and clean, sixty-odd games of consistent production, cleared every gate without friction. We rated him confidently. He played eleven minutes that night instead of his usual eighteen, because his team was managing a back-to-back and the coaching staff rotated his line more aggressively in the third period than in any of the preceding games we had on file. The output was roughly proportional to the minutes. The corpus was not wrong about the player. It was wrong about the context the player would be placed in, and we had no gate for that.
That is the version of the story where the corpus entry is technically correct and useless — the number described a real player doing real things, and still told us nothing about the night in question. We have written about that specific failure mode before, but the rotation and rest problem cuts deeper than a single bad entry. It is structural, and it is different in every sport we work in.
Discover the surprising reasons behind the things, rules, habits, and systems we encounter every day.
How Each Sport Punishes the Same Assumption Differently
In basketball, load management is visible and loud. A player listed as probable carries a different prior than a player who has played thirty-eight minutes for six straight games. The gate for availability is almost never the problem — the problem is that available is not the same as expected to feature, and in basketball those two states diverge most dramatically when a team is deep in a schedule cluster. A player can be available and still receive twenty-two minutes instead of thirty-four. The corpus records what he does with full deployment. The night in question may not be a full-deployment night.
Hockey compounds this because the rotation happens mid-game and is not announced. Line combinations shift between periods. A forward who averaged sixteen minutes over forty games may skate twenty minutes one night and ten the next, based on score state, penalty trouble, and decisions made in the second intermission. The corpus is blind to all of that in advance. We can see the historical average, but the variance around that average is itself a function of team context that changes game to game.
Soccer is worse in a specific way: the five-substitution window means a player can be removed at the fifty-fourth minute and the corpus entry for that game records less than an hour of contribution. String enough of those together and the average looks suppressed without any signal about why. Is this a player being managed through a heavy fixture schedule, or a player who is losing the coaching staff's confidence? The corpus cannot tell you. The gate for "expected to feature" becomes almost philosophical — a player can start a match and contribute for fifty minutes and the corpus will file that next to a full ninety-minute performance as if they are the same kind of observation.
Baseball is the outlier. Rotation in the pitching sense is predictable and published; rest for position players is more visible than in any other sport we track. The problem there is different — it is not that we cannot see the rest pattern, it is that the rest pattern interacts with handedness matchups, bullpen usage, and lineup construction in ways that make the prior-performance corpus genuinely difficult to condition on. The sample sizes are large enough that we trust the corpus more in baseball than anywhere else. We are also probably overconfident about that trust.
What We Built to Account for Minutes and Deployment
After the hockey episode, we spent about three weeks trying to build a minutes-adjustment layer into the corpus. The idea was simple: normalize each performance entry to a per-minute rate, then project forward using an estimated-minutes figure derived from recent deployment patterns. If a player had averaged sixteen minutes over his last twenty games but had been trending toward fourteen in back-to-back situations, the projection would use fourteen as the denominator rather than sixteen.
Riku, who handles most of our hockey corpus work, was skeptical from the start.
"You're estimating minutes using a sample of back-to-backs, which is maybe eight games, and then you're multiplying that estimate by a per-minute rate that has its own error. You haven't reduced the uncertainty. You've just moved it and given it a formula."
He was right, and we built it anyway, because the alternative — ignoring the problem — felt worse. The per-minute normalization worked reasonably well in basketball, where the minutes estimates were more stable and the corpus entries were numerous enough to produce reliable rates. In hockey it introduced noise we had not had before. In soccer it was nearly useless, because per-minute rates in soccer are not stable across different match states: a player who spends sixty minutes in a game where his team is chasing a deficit is doing something categorically different from a player who spends sixty minutes protecting a lead, and the corpus treats both as sixty-minute performances.
We also tried flagging schedule density explicitly — tagging any rating produced when a team was playing its third game in five days. The flag was meant to trigger a shrinkage adjustment, pulling the confident estimate closer to the base rate. That part worked in the sense that it reduced our overconfidence. It did not improve accuracy. It mostly made us less wrong in the way that saying nothing makes you less wrong.
Where the Adjustment Made Things Worse, Not Better
The cost was concentrated in a specific kind of player: the workhorse who actually does not get rested. We had conditioned the model to expect reduced deployment in high-density schedule windows, and some players simply do not get reduced deployment. Their coaches play them heavy regardless of the calendar. When we applied the rest-adjustment shrinkage to those players, we pulled their estimates toward a base rate they had spent years demonstrating they did not belong near. The corpus was right about them. The adjustment was wrong about the situation.
This is the version of overconfidence that is hardest to catch, because it looks like caution. We were shrinking estimates, which is supposed to be the conservative move. But shrinking in the wrong direction, away from a player who has earned his corpus through exactly the kind of durable heavy-minutes performance the adjustment was discounting, is not caution. It is a different kind of mistake wearing caution's clothes.
We noticed it when we ran the grading window and found that our worst-performing ratings that quarter were clustered among players with large, clean corpora who had been flagged for schedule density. The flag was functioning as a penalty on the most reliable data we had. Nadia pulled the breakdown and the pattern was unambiguous: the players we had trusted least, on the basis of the schedule adjustment, were the ones the corpus had been most consistently right about. A run of good results in the weeks before had made us think the adjustment was working. It was not. We had been grading ourselves on the wrong window.
What Stayed, What We Dropped, and What We Are Still Not Sure About
We kept the per-minute normalization in basketball, where the estimates are stable enough to be worth the added complexity. We dropped it in hockey and soccer, where the instability of the estimates compounded rather than absorbed the underlying uncertainty. The schedule-density flag stayed, but it no longer triggers automatic shrinkage — it triggers a review, which means a person looks at the player's specific rest-and-rotation history before any adjustment is applied. That is slower and less systematic, but it is less wrong.
The gate structure changed in one specific way: we added an explicit check for whether a player's recent deployment pattern has diverged from their historical average by more than a threshold we set at roughly fifteen percent. If it has, the candidate is flagged for manual review before scoring. This is a gate, not a model — it does not produce an adjusted estimate, it produces a pause. Most of the time the review confirms the corpus and we proceed. Occasionally it surfaces something the corpus would have missed: a player being quietly managed through a minor injury, a line combination that has shifted, a coach who has started shortening rotations in a way that is not yet reflected in the historical average. The gate does not fix the problem. It slows us down enough to notice it.
The part we are still not sure about is soccer. The five-substitution window and the match-state dependency of per-minute rates mean that the corpus entry for a soccer player is a noisier object than the same entry for a basketball or hockey player. We have been committed to the principle that sample size beats recency since the shop started, and that principle holds — but in soccer it holds less tightly than anywhere else, because the sample is partially composed of observations that are not the same kind of thing. A full ninety-minute performance and a fifty-four-minute performance before substitution are not the same observation. We file them together because we do not have a better structure. That bothers Riku more than it bothers me, which probably means he is right.
The honest version of what we learned is that the corpus is a record of what a player did under conditions that may not repeat. In basketball those conditions are stable enough that the record is useful most of the time. In hockey and soccer, the conditions are managed dynamically by coaches making decisions we cannot see in advance, and the record is useful until it is not — and we do not always know which situation we are in until the grades come back. Whether that uncertainty is a solvable problem or just the shape of the sport, I am genuinely not sure.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.