PlayerGem

A fictional shop writing about method. No picks, no lines to act on, and nothing here is betting or investment advice. What this is.

No Single Night May Dominate the Slate

There was a stretch — four or five weeks in the middle of a basketball season — when the shop's output was, by every internal measure, impressive. The hit rate was up. The calibration looked clean. We were citing the same three teams in roughly half of our active ratings, and nobody noticed, because the numbers were behaving.

Then those three teams hit a condensed scheduling block on the same night. They were all playing. We had more exposure to a single evening's variance than we had ever intended, and we had not intended any. The night went sideways in the way nights do — injuries, a postponement, a performance that was genuinely anomalous by any historical measure — and our calibration report for that window looked like something had been dropped from a height. It had been. We had dropped it.

The rule we wrote afterward is blunt: no single team, no single night, and no single statistical category may account for more than a defined share of the ratings active at any moment. We call it the concentration cap, and it is one of the few rules in the shop that has never been seriously challenged since we wrote it. That is either a sign that it is correct or a sign that the lesson was painful enough that nobody wants to revisit it. Possibly both.

Trading Strategy Mechanics Explained

Learn how trading strategies, execution, market regimes, and risk work—without signals or hype.

Learn more

How One Busy Night Broke Four Weeks of Good Work

The problem was not the bad night itself. Variance happens, and a shop that falls apart at the first anomalous result was not built to last. The problem was that we had let our confidence cluster without noticing it had clustered.

When we went back through the records, the pattern was obvious in retrospect. The three teams we kept citing were all playing in a high-pace style that our corpus happened to model well. They had long histories, clean data, and the kind of statistical consistency that makes a rating feel solid. We were not wrong to rate them highly — individually, each rating was defensible. The failure was structural: we had never asked what would happen if all of them were active on the same night and all of them went wrong simultaneously.

Mara, who runs our grading reviews, put it plainly in the post-mortem:

"Each one was fine. The problem was that 'each one is fine' is not the same as 'all of them together are fine.' We were thinking about ratings in isolation and grading them in isolation, and then we were surprised when they failed together."

This is a version of a mistake the shop has made in other forms. We once chased a pattern that turned out to be two data points dressed up as a trend, and the structure of that error was similar: local evidence that looked convincing, never tested against what would happen if the local evidence was wrong all at once.

Writing the Cap Before We Fully Understood What It Was Capping

The first version of the concentration cap was numerical and arbitrary. We said no single team could appear in more than a fixed share of active ratings at one time — we settled on a threshold after an afternoon of argument and then adjusted it twice in the first month. We also said no single night could account for more than a defined portion of the total corpus being evaluated in any given window. We did not have a rigorous justification for the exact numbers. We had a justification for the principle and we picked the numbers by feel, which is not something we usually admit.

The statistical category cap came later and was harder to define. The concern was that if we were rating players heavily on one kind of output — say, a peripheral counting stat that had been reliable all season — and that stat category turned out to be noisy in a way we had not modeled, we would be wrong in the same direction across many ratings simultaneously. Correlated errors are worse than independent errors by a factor that compounds quickly, and we had been treating our ratings as if they were independent when they were not.

The cap does not prevent us from rating players on the same team or the same night. It prevents any one of those groupings from becoming so large that a single bad outcome — an injury, a scheduling anomaly, a referee pattern we had not accounted for — could invalidate a disproportionate share of what we had produced. It is, in the language we use internally, a rule against believing ourselves too much in one direction at once.

This connects to something we think about often when reading a line as evidence about beliefs rather than as a verdict. A market that has concentrated heavily on one outcome is not necessarily wrong, but it is worth asking what would have to be true for it to be wrong in bulk. We ask the same question of ourselves.

What the Cap Took From Us, and Where We Applied It Wrong

The honest cost was apparent precision. When you force a distribution across teams, nights, and categories, you sometimes exclude a rating that is individually strong because the slot it would occupy is already full. We lost ratings we thought were good — and some of them probably were good — because the cap said the category was saturated.

That felt like waste. It still does, occasionally. There is something uncomfortable about excluding a well-supported rating for structural reasons rather than evidential ones. Colleagues have argued, not without justification, that a cap applied mechanically is its own form of overriding the data.

We also applied the cap incorrectly in at least one documented case. We were tracking a hockey player across a long stretch of games and had accumulated what we considered a genuinely robust corpus — more than enough to say something confident about a narrow range of output. The cap flagged the night as over-concentrated because we already had two other hockey ratings active. We excluded the third. The third would have been our most accurate rating of the month.

We recorded it. It is in the grading log. We do not know whether that means the cap was wrong in that instance or whether we got lucky with a rating we should not have trusted. This is, in miniature, the same problem we described when we wrote about the month our accuracy was highest and our calibration was worst — being right and being well-calibrated are not the same thing, and a single accurate exclusion tells you almost nothing about whether the rule is sound.

The cap has no mechanism for recognizing when it is costing us something real versus when it is correctly overriding a local confidence that would have embarrassed us. That is a design flaw we have not solved.

The Version of the Cap That Survived

We kept the cap. We adjusted the thresholds twice more after the hockey incident, and we added a review step: any rating excluded solely by the concentration rule is flagged for a post-hoc check, so we can track whether the exclusions are saving us or costing us over time. The aggregate is the only thing that matters. One exclusion that was wrong does not indict the rule, just as one exclusion that was right does not vindicate it.

The category cap is the part of the rule we are least certain about. Team concentration and night concentration are relatively easy to measure. Category concentration is harder because categories are not independent — a player who scores also tends to assist, and a player who assists tends to appear in games with high possession counts, and a player in games with high possession counts is generating data that correlates with other players in similar games. The categories bleed into each other, and our cap treats them as if they do not. We know this is a simplification and we have not found a clean way around it.

What the cap did change, permanently, is how we argue about individual ratings. Before it existed, the argument was almost always about the rating itself: is this corpus large enough, is the line reading correctly, is the shrinkage appropriate. Now there is a second layer: even if the rating is good, does it fit the current distribution. Mara calls this "the portfolio question," and she is right that it is a different kind of question than the ones we were trained to ask. It is less satisfying to answer, because the answer is sometimes "yes, it is good, and we are not using it."

The cap is also, we have come to believe, a check on a particular kind of overconfidence that is hard to catch any other way. When the shop finds a method that works — a corpus structure, a gate calibration, a way of reading a specific statistical category — the natural response is to apply it as widely as possible. The concentration cap is the rule that says: even if the method is working, do not let it fill the room. Sample size beats recency, and structural discipline beats a run of good results. That is the principle the cap is built on, and it is the principle we have to remind ourselves of most often.

What we have never fully settled is whether the cap is a rule about risk or a rule about epistemics — whether we are protecting the output or protecting ourselves from a specific kind of motivated reasoning. It might be both, and it might be that the distinction does not matter as much as we think it does.

Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.

Seven desks. The method, not the picks.

Start from the top