Research · pre-registered study
The midterms Brier skill curve, pre-registered
The Scopes Desk
How does a prediction market's accuracy compare with a poll-based estimate as Election Day approaches? This plan pre-registers one answer: the mean Brier score of the market and of the Scopes Fair Value at nine fixed horizons before November 3, 2026, scored the same way for both, with the method, the race universe, the sample gate, and the verdict lines all fixed here before the races resolve.
This is the horizon-by-horizon accuracy curve. The five separate claims about whether markets are downstream of polls, inherit their bias, lean by party, and add information were pre-registered earlier, on September 5, in Five claims about prediction markets and polls; they are not restated here. The companion contests scoring frontier models against the market are 2026 Senate control and the MLB postseason.
The specification
Universe
Every race on the polls roster that carries both a liquidity-gated market and a published Scopes Fair Value under the standing eligibility rule. The Senate-control and House-control markets are included as their own rows. A race that never cleared the composite gate (≥2 models, or ≥3 qualifying pollsters for the Poll Average) at a horizon is excluded at that horizon, not filled.
Horizons
90, 60, 45, 30, 21, 14, 7, 3 and 1 days before Election Day (November 3, 2026).
Probability used (pre-registered rule)
For each source at each horizon, the as-published probability from the last snapshot strictly before the horizon time: the market; the Scopes Fair Value; each named model where present; and the Scopes Poll Average where it was live. A composite is never recomputed with later methodology; the value the record holds is used, with its methodology_version. Named-model lines are limited to rights-clean, independent sources (open or display licensed, not a market-input model, and in the current composite slate), consistent with the standing source-rights policy; the headline market versus Scopes Fair Value comparison is unaffected.
Score
Brier per race per horizon per source against the resolved outcome (0/1); the mean across races per horizon; and the paired market-minus-composite difference with a cluster bootstrap by race (1,000 resamples) for the confidence band.
Sample gate
8 races with both sides available at a horizon. Below 8, the horizon's point is plotted with an early-data label and reads "insufficient power, no verdict." This gate was set from the availability audit below (roster of ~13 races; ~10 clear both sides at the mature horizons), not chosen after seeing results.
Output
One chart (x = days before Election Day, y = mean Brier, one line per source, shaded band on the paired difference) plus the table, and an expandable per-race table so any single race's curve is inspectable. Rendered server-side with the numbers in the HTML, in the dataviz house palette, light and dark.
Verdict templates (filled mechanically after resolution)
Chosen only by the sign and the significance of the paired difference at that horizon. The braces are filled from the computed numbers; no other text is added.
60 days out
At 60 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).
30 days out
At 30 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).
7 days out
At 7 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).
Data availability at registration
The collectors began on July 15, 2026 and the roster grew over the summer, so the earliest horizon is thin. We disclose that here rather than backfilling it. Races with both a market snapshot and a composite value within 24 hours before each past horizon:
| Horizon | Date | Races with both sides | Status |
|---|---|---|---|
| 90d | Aug 5 | 1 | Below the gate: no verdict |
| 60d | Sep 4 | 10 | Meets the gate |
| 45d | Sep 19 | 10 | Meets the gate |
| 30d | Oct 4 | — | Future at registration; fills at 6h cadence |
| 21d | Oct 13 | — | Future at registration; fills at 6h cadence |
| 14d | Oct 20 | — | Future at registration; fills at 6h cadence |
| 7d | Oct 27 | — | Future at registration; fills at 6h cadence |
| 3d | Oct 31 | — | Future at registration; fills at 6h cadence |
| 1d | Nov 2 | — | Future at registration; fills at 6h cadence |
Exclusions, pre-registered: governor races (no published composite: no forecast-model coverage and the Poll Average is below its pollster gate) and any Kalshi-only race whose composite is gated are excluded until and unless they clear the gate. Snapshot cadence for the market and the composite through Election Day is at least every 6 hours (ingest at 00, 06, 12, 18 and 07 UTC, with the composite recomputed each run).
The selection rule
Both sides come from the append-only record. The market side is price_snapshots (per contract, keyed to the race outcome); the composite and each model are benchmark_probs rows keyed by source (scopes-composite, scopes-poll-avg, statsheet-model). At each horizon each source's value is the last row strictly before the horizon time. The canonical availability query is pinned below; the scoring code and its methodology version are published with the results.
Horizon selection (per source): last snapshot strictly before the horizon
select implied_prob, captured_at from <price_snapshots | benchmark_probs> where <contract_id | outcome_id> = :id and captured_at < :horizon_time order by captured_at desc limit 1;
What this can't test
One cycle is a small sample. Where a horizon lacks the races for a stable estimate, its verdict reads "insufficient power, no verdict," by rule, rather than overclaiming on thin data. Brier measures calibration and resolution against a binary outcome; it does not adjudicate why a source was right or wrong. Those questions are the separate, already-registered claims above.
Follow the inputs as they fill in on the poll desk and the scorecard. Results here are computed after November 3, 2026.