Scopes

Research · methodology note · Sept 5, 2026

When our own benchmark lost to the market

The scorecard keeps score on our own fair value, not only the market. So when a review of the resolved record showed the market beating our benchmark on two desks, that is a result we publish, not one we bury. This note covers the energy and weather desks, both priced by a statistical fair value: what the record showed, why, and what we changed.

What the record showed

Across the resolved flags on both desks, the market (Kalshi) closed on the right side of our fair value more often than not:

DeskResolvedMarket wonFair value
Energy29964.5%realized-vol fair value
Weather7363.0%NWS-forecast fair value

Weather was worse than the headline suggests: its Brier score of 0.290 is above the 0.25 you get by printing 0.50 for every contract. The fair value was confidently wrong.

It was not timing

The obvious suspect is contamination near settlement, so we split the energy flags by how long before resolution each was created. The loss is not near the close. It sits at long lead, where a benchmark has to have a view:

Lead time at flagFlagsMarket won
more than 24h27965.6%
6 to 24h1650.0%
under 6h450.0%

The whole loss sits in the more-than-24h bucket (279 of 299 flags). A freeze window near settlement, the obvious fix, would have removed the two harmless buckets and kept the one doing the damage.

It was not stale data

We also added a timestamp for the age of each fair value’s input, so we could measure staleness instead of assuming it. The weather forecast was fresh: 1.4 hours old on average, none over six. Energy’s price anchor is fetched live at each run. Neither desk’s problem is old data.

It was the tails

Splitting energy by moneyness points straight at the cause. Near the money the fair value is fine. On far strikes, where it prints a confident probability near zero or one, the market wins seven times in ten:

StrikeFlagsMarket won
Far (deep out/in the money)22970.3%
Near the money (fair 0.35 to 0.65)7045.7%

Weather tells the same story from the other side. When we sort its forecasts into probability bands and check how often each happened, the ranking is inverted: the contracts it called least likely happened most, and the ones it called likely never did.

Forecast bandFlagsForecastHappened
0.00 to 0.06341.8%38.2%
0.27 to 0.301428.3%7.1%
0.65 to 0.671266.7%0.0%

Both are the same failure. A distribution with tails too thin puts too little probability on outcomes away from the center, so its most confident calls are the ones that lose. Energy used a normal (Gaussian) return model; weather used a forecast band that was too narrow.

What we changed

Reading it honestly

A benchmark you never check against the market is a claim, not a measurement. This is the scorecard the review now governs.