Research · methodology note · Sept 1, 2026
In-play contamination: what changed and why
The scorecard compares a market's price to an independent pregame benchmark: a de-vigged sharp book for games, an options- or realized-vol fair value for index and commodity contracts. That comparison is only meaningful while the benchmark still describes the same uncertainty the market does. Once a game is underway, or a daily contract is minutes from settlement, that stops being true.
The problem
The benchmark is a pregame line, it doesn't update once play begins. The market does. So a snapshot taken after kickoff isn't market-vs-fair-value; it's a live price against a frozen number. Near settlement the failure is the mirror image: the market and the fair value both collapse toward the outcome, so any remaining "gap" is convergence noise, not a measurable disagreement. Flags minted from either kind of snapshot don't measure what the scorecard claims to measure.
The rule change (flagging v2)
Each snapshot now carries a comparison_valid flag. It is false for in-play snapshots (sports, once the game has started) and near-expiry snapshots (stocks within 2h of settlement, energy within 24h). The scorecard counts comparison-valid flags only. Invalid snapshots are not deleted, they stay in the archive in full (the convergence window is valuable history), simply labeled and excluded from the score, ranking, and the hero. The record is append-only: prior flags remain, tagged by the rule that created them.
What it did to the record
The split at cutover: the share of resolved flags where the fair value proved right, under the old rule (all flags) versus the new one (valid-only):
| Desk | v1 (all) | v2 (valid) | resolved | excluded |
|---|---|---|---|---|
| Stocks | 44.4% | 63.2% | 36 → 19 | 17 |
| Energy | 27.7% | 25.8% | 256 → 236 | 20 |
| Sports | 49.1% | 49.1% | 226 → 226 | 0 |
| Weather | 37.5% | 37.5% | 56 → 56 | 0 |
| Rates | 100.0% | 100.0% | 1 → 1 | 0 |
Reading it honestly
- Sports, weather, rates were unchanged. Sports flags were already created pregame-only; the other desks have no in-play or near-expiry exposure. The entire correction lives in stocks and energy.
- Stocks: the correction runs in our favor. Nearly half the flags (17 of 36) were minted near settlement, and most of those resolved wrong: near-expiry noise. Removing them lifts the clean rate from 44.4% to 63.2%. The old number was understating the desk. (The valid sample is small, 19 flags, so read it as early.)
- Energy: it runs against us. The 20 excluded flags were "right" about half the time, correct by convergence, not by foresight, so removing them nudges the rate down, 27.7% to 25.8%. We report it anyway. The point of the rule is to strip contamination wherever it sits, including where it flattered us.
Every prior flag is still in the record, tagged with the rule that created it; nothing was rewritten. This is the scorecard the change now governs.