The Tennis Gate That Told Us Nothing About Tennis
For most of the first year we ran tennis through our gates, one of them had a near-perfect pass rate. The availability gate — the one that asks whether a player is actually expected to feature in the event being evaluated — cleared somewhere above ninety percent of candidates. We treated this as a sign that the gate was working. It was not. A gate that clears almost everything is not a gate. It is a formality with a checkbox.
The problem was structural and embarrassingly obvious once we found it. In team sports, availability is genuinely uncertain: a player can be scratched the morning of a game, held out of a rotation, moved off a lineup for reasons that are never fully explained. The gate exists because those absences are unpredictable and the corpus cannot account for them. In tennis, a player who has entered a draw and not withdrawn is, by definition, available. They are on the schedule. The schedule is public. We were confirming something the tournament website already confirmed, then logging it as a passed gate and moving on feeling diligent.
What we were not doing — and this is the part that cost us — was asking the question the gate was supposed to answer in the first place: is this player in a condition to perform at the level the corpus describes? Those are very different questions. One is about presence. The other is about readiness. We had built a gate for presence and convinced ourselves it was measuring readiness, and tennis punished us for that distinction in a way that basketball and hockey simply had not.
A free online course in data analytics with Python: statistics, visualization and finding the signal in the numbers.
Why tennis has a readiness problem that the availability gate was never designed to catch
Tennis is an individual sport played on a continuous draw structure. A player who wins on Monday plays again on Wednesday. A player who played three sets on Monday and five sets on Wednesday is not the same player they were at the start of the week, and the corpus — built on their full performance history — does not know any of that happened. The corpus sees a body of work. It does not see a body that played ten sets in four days.
In team sports, the schedule creates natural rest. Even a player who features heavily has teammates absorbing minutes, and the game structure limits individual load in ways the corpus implicitly captures. Tennis has none of that buffering. The load accumulates within a single event, across rounds, and a player can go from fresh to visibly compromised between the time a draw is set and the time a match is played.
We knew this in the abstract. We had discussed it. Renata, who handles tennis in the corpus, had flagged it twice in review meetings. The problem was that knowing something and building a gate around it are different operations, and we had done the first without doing the second. The availability gate cleared the player because the player was in the draw. The readiness question sat in the notes from Renata's second meeting and never became a disqualification criterion.
"We had a gate for 'is this person showing up.' We needed a gate for 'is this person intact.' Those are not the same sport."
— Renata
This is a version of a problem we have written about before: the tendency to carry a gate built for one sport's assumptions into a sport that doesn't share them. In that case the transfer was explicit. Here it was passive — we never consciously decided to apply the team-sport availability logic to tennis. We just never decided not to.
What we tried when we finally built a gate that asked the right question
The replacement gate had two components. The first was a round-load counter: how many matches had the player completed in the current event, and what was the cumulative set count across those matches. We set a threshold — not published here because a threshold without context is just a number — above which a candidate would be flagged rather than cleared, and the flagged candidates would require a second review before scoring.
The second component was a rest-interval check. If fewer than a certain number of hours separated a player's last completed match from the scheduled start of the match being evaluated, the candidate was flagged regardless of round count. This was meant to catch the scenarios where a late finish in one round compressed the recovery window before the next.
Both components required data we did not already have in the corpus. We had performance records. We did not have match-end timestamps for historical draws, which meant we could build the gate prospectively but could not backfill it. That gap mattered: the calibration work we did after building the gate had no historical baseline to compare against, because the old gate had been logging passes where the new gate would have logged flags. We were starting a new record, not correcting an old one.
The grading implications of that gap are something we are still working through. Grading windows are already fragile enough without introducing a structural break in the gate logic partway through a season. We made the change anyway because the old gate was clearly wrong, but we want to be honest that the first year of the new gate's calibration history is not directly comparable to anything that came before it.
What the passing gate actually cost us in the corpus and why we missed it for so long
The direct cost was a corpus that overrepresented late-round performance from players who had accumulated heavy loads. Because the availability gate cleared those candidates, their late-round matches entered the scoring pool. Late-round matches in a draw are, structurally, between the most successful players in the event — which means they are a biased sample of performance, drawn from a specific physical and psychological context that the corpus treated as ordinary.
We did not notice this for a long time because the inflation was symmetric in a way that hid it. If every player's corpus was slightly inflated by late-round data, the relative rankings were not obviously distorted. The ratings looked internally consistent. What they were not was accurate in absolute terms — the corpus was describing players as capable of a level of output that they could sustain in the middle rounds of a fresh draw, not across a full event's worth of accumulated load.
The reason we missed it for so long is less flattering. Tennis was the most recently added sport in our scope, and the people who built the initial gate structure — myself included — imported the logic from the sports we knew better without stress-testing whether it transferred. There is a version of this failure that looks like carelessness, and I want to be honest that it partly was. We were moving quickly and the gate seemed to be working because it was producing a number. A gate that produces a number is not necessarily a gate that is measuring the right thing. Inherited gate logic has a way of persisting past the point where anyone remembers why it was written that way in the first place.
Renata was right, and she said so twice, and we did not act on it quickly enough. That is the error. It is in the record.
What we kept from the old gate and what the new one still cannot see
We kept the withdrawal check. A player who has formally withdrawn from a draw is a genuine disqualification, and the original gate caught those correctly. We also kept the draw-entry confirmation as a first pass — it is still useful as a precondition, even though it is not sufficient on its own. The new gate runs the withdrawal check first, confirms draw entry, then runs the load and rest-interval components. The old logic became step one of a longer sequence rather than the whole sequence.
What the new gate still cannot see is the qualitative condition that experienced tennis observers read without difficulty: a player who is moving differently, protecting a shoulder, or visibly rationing effort in a match they are winning comfortably. That information exists. It lives in match footage and in the observations of people who watch the sport closely. We do not have a systematic way to incorporate it, and the attempts we made to formalize it produced features that were noisy enough to reduce accuracy rather than improve it.
This is the part of the problem where instinct holds an advantage the corpus cannot match. A scout who has watched a player across a full event carries a readiness signal that our gate approximates with load counts and timestamps. The approximation is better than nothing. It is not better than the scout, and we do not think it is. The shop's whole position on this — and it has not changed since we started — is that systematic measurement earns its place alongside instinct, not above it.
We are also aware that the round-load threshold we set is, itself, a number we chose without a strong empirical basis. We had enough data to make a reasonable argument for it. We did not have enough data to be confident in it, and the distinction matters. It is the kind of methodological choice that looks clean in a write-up and feels much less clean when you are sitting with the calibration output six months later wondering whether you set it one round too high.
The question that has stayed with us is whether a readiness gate for tennis is fundamentally buildable from the outside — whether load counts and rest intervals are a proxy for something that can only really be seen, and whether the gap between the proxy and the thing will always be large enough to matter. We do not know. We built the best gate we could from the data we had access to, and we are watching the calibration numbers with less confidence than we usually admit to.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.