Most Candidates Should Die Before You Score Them
People assume the interesting part of an evaluation shop is the scoring. It is the part with the mathematics in it, so it is the part that sounds like the work.
In our system the scoring is perhaps a fifth of the code and, on a normal day, decides almost nothing. The overwhelming majority of what happens is candidates being thrown out before anything is calculated at all.
Analytics, observability, and AI-driven insight from test runs.
We Built the Scoring First and the Rejections Last
We got the order exactly backwards, and I think most people would. We built a way to evaluate a candidate, then built the machinery to feed candidates into it, and only later — after several months of results that were somehow both plausible and useless — started asking which candidates should have been in the queue at all.
The symptom was that our output was full of things that were technically correct and practically meaningless. A rating on somebody whose availability was unknown. A rating on an event that had already begun. A rating on a player whose record could not be matched to an identity with any confidence, so the history we scored belonged to somebody else with a similar name.
None of those are scoring failures. The scoring did precisely what it was asked. They are failures of admission, and we did not have an admissions process — we had a queue.
Inverting It: Everything Is Rejected Until It Is Not
The rebuild flipped the default. Nothing is eligible. A candidate has to survive a series of independent disqualifications, each one narrow enough to be tested on its own, before the scoring is permitted to look at it.
The gates are dull individually and that is the point. Is there enough history. Is the player expected to feature. Can this candidate be linked to a specific future event. Has that event already started. Does the identity resolve to exactly one player. Is the sport currently in season, or are we looking at something synthetic left over from a prior one.
On a typical day the great majority of what enters the pipeline never reaches the scoring. That ratio bothered us at first — it looks like waste. It is not waste. It is the system declining to make claims it cannot support, which is the only reason any of the claims it does make are worth anything.
“Every one of these gates exists because something embarrassing got through,” Ellen said, going through the list. “There isn't a single one we thought of in advance.”
What Aggressive Rejection Costs Us
It costs volume, and volume is not nothing. There are days when almost nothing survives, and a shop with an empty output has to resist the urge to relax something.
We have relaxed something, twice. Both times it was late in a quiet week and both times the reasoning was that a particular gate was probably too strict. Both times we produced ratings we later graded badly and had to keep on the record. The second occasion is the more embarrassing because we had the first one written down.
The other cost is real and harder to price: we are certainly rejecting good candidates. A gate that catches every bad case will catch some acceptable ones, and we have no way to count those, because a rejected candidate is never rated and therefore never graded. Our error rate only measures the things we let through. That asymmetry is baked into the design and I do not know how to remove it.
One Gate, One Reason, One Test
What we kept is a structural rule rather than any individual gate: every gate is a single condition, with a written reason, and a test that fails if the gate stops rejecting what it was built to reject.
That last part is the one that has saved us most often. A gate that silently stops working looks exactly like a gate that has nothing to do, and without a test you cannot tell the difference until something bad reaches the output. We learned that the slow way.
We also log every rejection with the gate that caused it. Not for analysis, particularly — for the much simpler reason that when the output looks strange, the first useful question is which gate suddenly changed its mind.
The scoring has been rewritten four times. The gates have only ever been added to. I think that ratio says something, though I am not certain what.
Note: PlayerGem is a fictional analytics shop and these accounts are invented. Nothing here is a pick, a recommendation, or betting or investment advice, and the players, teams and competitions described do not exist.