Research · methodology note · Sept 5, 2026
When our own benchmark lost to the market
The scorecard keeps score on our own fair value, not only the market. So when a review of the resolved record showed the market beating our benchmark on two desks, that is a result we publish, not one we bury. This note covers the energy and weather desks, both priced by a statistical fair value: what the record showed, why, and what we changed.
What the record showed
Across the resolved flags on both desks, the market (Kalshi) closed on the right side of our fair value more often than not:
| Desk | Resolved | Market won | Fair value |
|---|---|---|---|
| Energy | 299 | 64.5% | realized-vol fair value |
| Weather | 73 | 63.0% | NWS-forecast fair value |
Weather was worse than the headline suggests: its Brier score of 0.290 is above the 0.25 you get by printing 0.50 for every contract. The fair value was confidently wrong.
It was not timing
The obvious suspect is contamination near settlement, so we split the energy flags by how long before resolution each was created. The loss is not near the close. It sits at long lead, where a benchmark has to have a view:
| Lead time at flag | Flags | Market won |
|---|---|---|
| more than 24h | 279 | 65.6% |
| 6 to 24h | 16 | 50.0% |
| under 6h | 4 | 50.0% |
The whole loss sits in the more-than-24h bucket (279 of 299 flags). A freeze window near settlement, the obvious fix, would have removed the two harmless buckets and kept the one doing the damage.
It was not stale data
We also added a timestamp for the age of each fair value’s input, so we could measure staleness instead of assuming it. The weather forecast was fresh: 1.4 hours old on average, none over six. Energy’s price anchor is fetched live at each run. Neither desk’s problem is old data.
It was the tails
Splitting energy by moneyness points straight at the cause. Near the money the fair value is fine. On far strikes, where it prints a confident probability near zero or one, the market wins seven times in ten:
| Strike | Flags | Market won |
|---|---|---|
| Far (deep out/in the money) | 229 | 70.3% |
| Near the money (fair 0.35 to 0.65) | 70 | 45.7% |
Weather tells the same story from the other side. When we sort its forecasts into probability bands and check how often each happened, the ranking is inverted: the contracts it called least likely happened most, and the ones it called likely never did.
| Forecast band | Flags | Forecast | Happened |
|---|---|---|---|
| 0.00 to 0.06 | 34 | 1.8% | 38.2% |
| 0.27 to 0.30 | 14 | 28.3% | 7.1% |
| 0.65 to 0.67 | 12 | 66.7% | 0.0% |
Both are the same failure. A distribution with tails too thin puts too little probability on outcomes away from the center, so its most confident calls are the ones that lose. Energy used a normal (Gaussian) return model; weather used a forecast band that was too narrow.
What we changed
- Energy, now. The desk only flags near-the-money contracts, where the fair value has an even chance, instead of the far strikes where it was reliably wrong. This makes the desk honest and neutral rather than losing.
- Energy, the real fix. The return model moved from a normal distribution to a fatter-tailed one (Student-t), so far-strike probabilities pull back toward the middle instead of overcommitting. As those flags resolve, we can widen the near-the-money restriction back out on evidence.
- Weather stays dark. The desk was already unpublished; it stays that way until its forecast band is wider and corrected for any center bias, and its reliability curve is monotone. The table above is the data we fit against.
Reading it honestly
- The near-the-money energy sample is small (70 flags) and its win rate is a coin flip within error. The fix stops a loss; it does not yet prove an edge.
- The weather bands hold few flags each, so read the shape of the inversion, not any single cell.
- Nothing was deleted. Every prior flag stays in the record, tagged with the model version that created it; the fixes apply going forward.
A benchmark you never check against the market is a claim, not a measurement. This is the scorecard the review now governs.