Available Is Not the Same as Expected to Feature
The gate we installed first was the obvious one: is the player available? Injured reserve, suspension, illness — anything that formally removes a player from consideration. That gate is easy to build and satisfying to run because it has a clean answer. We ran it for about three months before we noticed it was doing almost nothing for us. The candidates who cleared it were still producing ratings that graded out wrong at an embarrassing rate, and the failure wasn't random. It had a shape.
The shape was this: the players who broke our grades weren't unavailable. They were available in the technical sense — dressed, present, on the active list — and then they played four minutes, or came off the bench in a blowout for a single shift, or were held out of a rotation we hadn't been told was changing. They were available. They were not expected to feature. We had been treating those two things as synonyms, and they are not even close.
This is the piece about what it took us too long to learn, and about the secondary gate we had to build once we admitted the first one wasn't doing the job we thought it was.
Victor Draemont’s notes on discipline, judgment, power, and playing the long game.
The Difference Between "On the Roster" and "In the Plan"
A player who is formally available sits somewhere on a spectrum that runs from "starts and plays full minutes" to "dresses and does not leave the bench." For most of our rating work, the useful part of that spectrum is narrow. We need candidates who are expected to feature — to play enough that the performance category we are measuring has a reasonable chance of being reached.
The problem is that "expected to feature" is not a field in any dataset we have ever seen. It is an inference. You build it from recent lineup patterns, from reported practice participation, from whether the team is in a scheduling stretch that typically compresses rotations, from whether a player returning from a minor injury has been described as "on a minutes restriction" or "full go." None of those sources are reliable on their own. Several of them contradict each other routinely. What disagreement between sources tells you is usually something useful about uncertainty, but at this gate it mostly told us we needed a higher threshold before we'd let a candidate through.
The gate we eventually built had three components. First, a minimum recent-activity requirement — if a player had not logged meaningful time in a configurable window of prior appearances, they did not pass, regardless of their formal status. Second, a lineup-signal check: we required at least one corroborating indicator that the player was in the expected starting group or rotation, not just available as a reserve option. Third, a restriction flag: any reporting of a limitations qualifier — "game-time decision," "minutes cap," "easing back in" — triggered a hold, not an automatic disqualification but a pause that required a second read before anything was scored.
How We Tried to Automate the "Expected to Feature" Judgment
The naive version of this gate was a minutes-played floor. If a player had averaged above a threshold number of minutes across their last several appearances, they passed. It was simple, it was fast, and it caught some of the worst offenders — players who had been playing reduced roles for weeks and were still clearing our availability check on name recognition alone.
We ran that version for about six weeks. It improved things, and we said so internally. Mara, who handles most of our basketball corpus work, was the one who pointed out that we were measuring the wrong lag. "You're looking at whether they played recently," she said. "You're not looking at whether anything changed between the last game and this one." She was right. A player could average forty minutes a game for two months and then get a quiet rotation adjustment two days before a game we were rating. The minutes floor told us nothing about that.
"The availability check is a historical question. The featuring check is a forward-looking one. We kept building historical answers to a forward-looking problem."
— Mara
So we added the corroborating-indicator layer. We pulled from practice reports, from pre-game notes, from any structured source that described expected lineups rather than past ones. This was slower and dirtier than the minutes floor. The sources were inconsistent in format, inconsistent in timing, and occasionally just wrong. But it caught a category of candidate the minutes floor missed entirely: the player who had been featuring heavily and was quietly being managed down in the days before a specific game.
The connection to the maintenance work that makes the corpus usable is direct here. A gate is only as good as the data it reads. If the corpus doesn't flag that a player's recent-game participation is trending down, the gate never fires. We found several cases where our corpus had clean historical records and genuinely stale recency signals, and the gate passed candidates it should have held.
The Month We Graded Ourselves on the Wrong Population
Here is the thing we got wrong, and I want to be precise about it because the error was structural, not incidental.
When we introduced the new gate, we recalibrated our expected pass rate — the share of candidates we expected to clear the gate and proceed to scoring. The new gate was stricter, so fewer candidates passed, and our hit rate on the ones that did looked better. We read that as the gate working. It was not the gate working. It was the gate changing the population we were grading ourselves against.
By removing the ambiguous candidates — the players who might feature or might not — we had made our remaining pool easier to be right about. The harder cases had been filtered out, and we were congratulating ourselves on accuracy that was partly a function of having stopped attempting the hard cases. We have written about this pattern before in a different context, but we walked into it again here, which is the kind of thing that happens when you are watching your hit rate instead of your calibration.
The actual cost was about a month of grades that looked better than they were, followed by a recalibration session that was uncomfortable. We had to go back through the held candidates — the ones the gate had stopped — and ask whether any of them should have been scored and weren't. Some of them should have been. The gate was too aggressive on the restriction-flag component; anything with a qualifier was being held, and some of those qualifiers were noise. "Day-to-day" on a minor soft-tissue issue three days before a game, for a player with a history of playing through exactly that kind of thing, probably should not have been treated the same as "minutes restriction" for a player returning from a significant absence.
We had built one flag where we needed at least two, and we had spent a month not knowing it.
What the Gate Looks Like Now, and What It Still Gets Wrong
The version we run today has four components instead of three. The minutes floor stayed. The corroborating-indicator check stayed. The restriction flag was split into a hard hold — for explicit minutes caps and confirmed limited-duty designations — and a soft flag, which requires a second pass but does not automatically stop the candidate. The fourth component is a recency-weight check on the corpus itself: if the player's corpus record hasn't been updated within a configurable window, the gate holds regardless of what the activity data says, because stale corpus records are where questions about adequate history tend to hide.
This version is slower and requires more manual review than the original. The soft-flag category in particular generates a queue that someone has to actually read, which is not how we wanted to spend time when we built the first version. But the alternative — automating a judgment that genuinely requires context — produced the bad month described above, and we are not in a hurry to repeat it.
What it still gets wrong: late-breaking information. A rotation change announced two hours before a game, a player who was full-go in morning skate and then scratched for reasons that don't appear in any structured source until after the fact — these still get through. The gate is a pre-game instrument, and some of the most important featuring information arrives after we have already run it. We have not solved that. We have accepted it as a boundary condition and try to be honest about what our grades mean in light of it.
Theo, who does most of our gate-review work, has a phrase for candidates that clear all four components but still feel uncertain: "technically clear." It is not a formal category. It is a note in the review log. But it shows up often enough that it probably should be.
The question we keep returning to is whether "expected to feature" is even a stable concept — whether it describes something a player has, or something a coaching staff decides, or something that only exists in retrospect once you know how the game unfolded. We built a gate around it because we had to put something there. Whether what we built corresponds to the thing we were trying to measure is a question we can't fully answer from inside the process.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.