PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

What Baseball Taught Us About Our Hockey Gates

We built the hockey gates in a hurry. That is not an excuse; it is a cause. In the autumn of our second year covering the sport, we were still running a gate structure we had designed for baseball, and we were proud of it. The baseball gates had taken eighteen months to tune. They caught scratches, late lineup changes, players returning from minor-league assignments with no announced role. They were, by our own calibration record, the tightest we had ever written. So when hockey came onto the desk, we did what careful people do when they are in a hurry: we borrowed.

The logic seemed sound at the surface. Both sports play long regular seasons. Both have daily or near-daily scheduling windows. Both produce per-game counting stats granular enough to build a corpus around. We assumed the gates would need cosmetic adjustment — swap "batting order confirmed" for "line combinations confirmed," change a few field names, update the timing thresholds. We assumed this for longer than we should have.

What followed was not a catastrophe. It was quieter than that: a slow accumulation of candidates passing gates they should not have cleared, ratings issued on players whose actual role on a given night bore no resemblance to what the corpus described. We did not notice for almost six weeks, which is its own kind of embarrassment.

The stories behind the things around us

The origins and reasoning behind familiar things.

Read Why This Exists

The Gate That Baseball Built and Hockey Broke

In baseball, the availability gate is essentially a binary. A player is in the lineup or he is not, and by a predictable window before first pitch, that information is public and stable. Lineups are filed. They are rarely altered after posting. The gate we built asked one question — is this player confirmed to appear, and in what position in the order? — and it asked that question once, roughly ninety minutes before game time. In two seasons of baseball, that gate misfired on us fewer than a dozen times.

Hockey does not work this way. Line combinations are announced, adjusted, announced again, and sometimes abandoned in warmups. A player can be listed on the first line at morning skate and be a healthy scratch by puck drop, with no formal update anywhere we could reliably monitor. More importantly, a player's role within a confirmed lineup — power-play time, deployment against certain opposition, whether a coach was experimenting with a line shuffle — varied in ways that the binary available/unavailable gate simply could not capture.

The baseball gate asked: is he in? The hockey question we actually needed to ask was: in what capacity, and how stable is that capacity? Those are different questions. We were asking the wrong one for six weeks and issuing ratings against a corpus that described a player's full-deployment performance while the gate had cleared him for a night when he might see half his usual ice time.

Rebuilding the Gate Around Role Stability, Not Just Presence

The fix we designed was a two-stage availability gate. The first stage matched baseball's: confirmed to dress, no injury designation, not listed as a healthy scratch in any of the three sources we monitored. That part we kept almost verbatim. The second stage was new, and it was harder to build.

We defined something we called a role confidence window. For each candidate, we looked at the previous fourteen days of confirmed line combinations — not games played, but announced combinations — and measured how often the player had appeared in the same deployment context: same line, same power-play unit presence or absence, same rough usage tier. If that figure fell below a threshold we set through calibration rather than intuition, the candidate was held. Not disqualified — held, pending the warmup report.

The warmup report gate was the piece we were least sure about. We gave ourselves a hard cutoff: if confirmed line information from warmups was not available within our processing window, the candidate did not advance regardless of what the morning skate had shown. Riku, who handles most of the hockey corpus maintenance, put it plainly when we were designing the rule.

"Morning skate tells you intent. Warmups tell you decision. We were grading on intent and calling it information."

That distinction — intent versus decision — became the conceptual spine of the revised gate. Baseball had never forced us to make it because in baseball, the lineup card is the decision. In hockey, it is closer to a draft.

What the Borrowed Frame Actually Cost Us

The six weeks of misapplied gates produced ratings we have since gone back and marked. Of the candidates who cleared the old gate and received a rating during that window, a meaningful fraction — we counted thirty-one, though the exact number depends on how you define a deployment mismatch — were evaluated against corpus data that described a different version of their role than the one they actually played that night.

Some of those ratings survived anyway, which is not reassuring. A rating that is right for the wrong reason is a calibration problem waiting to surface later. What we found when we ran the grading pass was not a clean disaster but a smeared one: the hit rate looked acceptable, but the confidence intervals were wrong. We had been more certain than we had any right to be, and the outcomes had not yet punished us for it. They would have, eventually. They usually do.

The deeper cost was to the corpus itself. Because we had been issuing ratings on mischaracterized deployments, some of those outcomes had been folded back into the corpus as if they were normal full-deployment performances. We had to audit and reweight approximately four months of hockey data. That took longer than building the new gate.

I want to be precise about where the error lived: it was not in borrowing from baseball. Cross-sport transfer is one of the more productive things we do. The error was in borrowing the gate without first asking what the gate was actually measuring, and whether that thing existed in the new sport. We checked that the structure fit. We did not check that the underlying assumption held.

What Stayed, What Changed, and One Surprise

The first stage of the baseball gate — the binary availability check — stayed almost intact, because the question it asks is genuinely shared. Both sports require a player to be physically present and expected to play. That condition is sport-agnostic and the gate that checks it is portable.

The timing logic also transferred, with adjustment. Baseball's ninety-minute window became a two-window system in hockey: an early check at morning skate for initial corpus filtering, and a hard gate at warmup confirmation for final clearance. The principle — do not advance a candidate until the information is as stable as it is going to get — was the same. The execution had to account for the sport's different relationship with finality.

The surprise was the shrinkage interaction. When we tightened the role confidence threshold and more candidates were held at the second gate, the pool of rated players on any given night shrank. We had expected this. What we had not expected was that shrinkage — the statistical pull of confident estimates back toward the base rate — behaved differently on the smaller pool. With fewer candidates advancing, each rating carried more weight in the nightly distribution, and our aggressive shrinkage parameters, calibrated for baseball's larger nightly pools, were pulling estimates too hard. We had to recalibrate shrinkage from scratch for hockey, using hockey data only.

Riku's note on this, from the calibration review, was the one I saved: "Baseball taught us what a gate is for. Hockey taught us that the gate is not the whole method." That is probably the most honest summary of the project.

We still borrow across sports — it would be wasteful not to. But there is a question we now ask before any structure moves from one sport to another: not does this fit, but what does this assume, and does that assumption survive the crossing? We do not always know the answer before we try. I am not sure there is a way to always know.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top