Why We Refuse Anything Already Underway
The first time we violated the in-progress rule, we did it knowingly. A hockey candidate had cleared every other gate — corpus large enough, activity recent enough, availability confirmed — and then, about twenty minutes before we published, the first period ended and the game was already underway. We published anyway. The reasoning was thin: the candidate's stat in question was a second- and third-period phenomenon, so the first period's result was noise, not signal. We told ourselves the gate was a heuristic, not a law, and that this particular situation was an exception the heuristic hadn't anticipated.
It was not an exception. It was exactly the situation the gate was designed for, and we had just talked ourselves out of the gate's entire purpose in about four minutes of motivated reasoning. The rating we published was not graded against anything the market had believed before the game began — it was graded against a market that had already absorbed a period of live information we had not priced. The comparison was meaningless. We had measured the gap between our estimate and a belief that no longer existed.
We wrote that one up. It sits in the misses folder, and it is not the only entry there with the same cause. The in-progress gate is one of the few rules in our process that has never been relaxed, not because we are disciplined, but because every time we tested it the result was the same: a corrupted grade that we could not honestly count in either direction.
Discover concerts, sports, theater, festivals, comedy, and more with DanteVibes.
What "Already Underway" Actually Breaks
The shop's method depends on reading a line as a statement of pre-event belief. That is the whole premise: whoever set the line had to commit to a number before anything happened, and the interesting analytical question is what that commitment reveals about what they expected. Once the event starts, that commitment is no longer the one being tested. The line moves. The belief updates. The original statement is gone.
When we score a candidate against a line that has already shifted in response to live play, we are not measuring what we think we are measuring. We are measuring the gap between our estimate and a revised belief — one that incorporates information we did not have when we formed our estimate, and one that the original setter might not even recognize as their own number anymore. We have written about missing line movement before, and the in-progress case is the most severe version of that problem: the movement is not a signal we failed to read, it is a signal that arrived after our reading was already locked.
There is a secondary problem that took us longer to name. When a game is underway, the candidate's stat line is partially observable. Anyone looking at a basketball candidate's assist count at halftime already knows the first-half number. That partial observability changes how people think about the remaining total — and it changes it in ways that are not uniform, not linear, and not something our corpus was built to model. Our corpus is built on pre-game distributions. Injecting a mid-game candidate into that framework is like using a weather model trained on morning forecasts to predict whether it will rain in the next ten minutes. The inputs are structurally different from what the model was fitted on.
The Workaround We Tried and Why It Seemed Reasonable
For a period of about six weeks, we tried to salvage in-progress candidates by building a separate adjustment layer. The idea was to subtract the already-completed portion of the stat from the line, re-center the distribution on what remained, and treat the residual as a fresh candidate. If a baseball candidate needed twelve strikeouts across a full game and had four through three innings, we would reframe it as an eight-strikeout question over the remaining six innings and run that through the standard gates.
It was not an unreasonable idea on its face. Colleagues at other shops do something similar with live markets, and there is published work on in-game projection that takes roughly this approach. Renata, who handles most of our baseball corpus, thought the residual method was worth a proper trial. Her argument was that we were leaving a large class of candidates unscored for a rule that might be more conservative than it needed to be.
"The gate made sense when we had no adjustment. Once we had the residual framing, the original reason for the gate got weaker. I thought we should at least find out whether the grades held up."
We ran the residual method for those six weeks across baseball and basketball candidates. We documented the threshold conditions — minimum innings or periods remaining, maximum partial completion, minimum corpus size for the residual distribution — written down before we looked at the grades, which was the right instinct. We were trying to be careful. We were not careful enough.
What the Residual Method Got Wrong and What It Cost Us
The grades from the residual method looked acceptable for the first three weeks. Then we ran calibration and found something we should have anticipated: the residual candidates were systematically overconfident in the later stages of games. When a candidate was more than sixty percent complete, our stated confidence was running about twelve points above our observed accuracy. We were saying we were right at a rate we were not achieving. Calibration failures of that shape are the ones that matter most to us, because they mean the shop is not just wrong — it is wrong in a way it cannot see.
The reason, once we found it, was embarrassing in its obviousness. The residual distribution assumed that the remaining portion of a game was statistically independent of the completed portion. It is not. A baseball pitcher who has thrown ninety pitches through five innings is not drawing from the same distribution for innings six and seven as a pitcher who starts fresh. A basketball player who has two fouls at halftime is playing differently in the second half than our corpus assumed. The adjustment layer we built corrected for the quantity already accumulated. It did not correct for the state of the player, the game, or the opposition that the completed portion had already created.
We also found a grading problem we had not anticipated. Some of the residual candidates were graded against lines that had moved substantially during the game, which meant we were back to the original issue in a different form. We had built a method to handle partial completion and accidentally reintroduced the corrupted-comparison problem through the back door. Renata's summary of the six-week trial was brief: "We made it more complicated and got the same failure." That is in the misses folder too, alongside a rating from that same period that we should have caught at this gate and did not.
The cost was six weeks of analysis time, a calibration gap we had to explain in the next quarterly review, and a minor but real erosion of confidence in the gates generally — because we had spent six weeks treating one gate as negotiable, and that habit is hard to fully unlearn. Some of our other doctrine failures started the same way: one reasonable-sounding exception, held open just a little too long.
The Gate as It Stands Now, and the One Adjustment We Did Keep
The in-progress gate was restored to its original form: if the event has begun, the candidate is refused. No residual framing, no partial-completion adjustment, no sport-specific carve-outs. The gate now triggers on confirmation of first action — tip-off, first pitch, puck drop, opening whistle — and anything that arrives after that timestamp is dropped without scoring.
We did keep one thing from the six-week trial, though it is procedural rather than methodological. We now log every candidate that dies at the in-progress gate, with the timestamp of when it arrived and how far into the event it was. We did not do this before. The log has turned out to be useful not for salvaging in-progress candidates but for diagnosing where our intake process is slow. If candidates are consistently arriving after the gate has closed, the problem is upstream — something in how we pull the corpus, confirm availability, or move a candidate through the earlier gates is taking longer than it should.
The log has also given us a clearer picture of which sports generate the most late arrivals. Soccer and hockey produce the fewest; the events are long and our confirmation process tends to resolve well before kickoff or puck drop. Basketball produces the most, mostly because the gap between availability confirmation and tip-off is narrow and our baseball-trained instincts about timing do not transfer cleanly. That is a structural finding about our process, not about basketball itself — though it is consistent with what we have noticed more broadly about how sport-specific timing assumptions contaminate gates built for a different sport.
The gate itself is not under review. The trial settled the question, at some cost, and the cost is what makes the answer stable. If we had dropped the residual method because it was theoretically unsound, we would probably revisit it the next time someone made a theoretically sound argument for it. We dropped it because we ran it, graded it, and found the calibration gap. That kind of evidence is harder to argue with than a principle, and we prefer it.
What I am less certain about is whether the gate is catching the right thing for the right reason. We kept it because the residual method failed calibration — but calibration over six weeks in two sports is not a large sample, and a different adjustment architecture might produce a different result. The gate is correct, as far as we can tell. Whether it is correct for the reason we think it is correct is a question the current evidence does not fully close.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.