Scopes

Research · pre-registered study

The midterms Brier skill curve, pre-registered

The Scopes Desk

How does a prediction market's accuracy compare with a poll-based estimate as Election Day approaches? This plan pre-registers one answer: the mean Brier score of the market and of the Scopes Fair Value at nine fixed horizons before November 3, 2026, scored the same way for both, with the method, the race universe, the sample gate, and the verdict lines all fixed here before the races resolve.

This is the horizon-by-horizon accuracy curve. The five separate claims about whether markets are downstream of polls, inherit their bias, lean by party, and add information were pre-registered earlier, on September 5, in Five claims about prediction markets and polls; they are not restated here. The companion contests scoring frontier models against the market are 2026 Senate control and the MLB postseason.

Pre-registered September 30, 2026. The universe, horizons, scoring rule, sample gate, exclusions, and verdict templates below are fixed as of this date. Results are computed mechanically from the tracked record after every roster race resolves, with no discretionary analysis and no re-specification. The plan is registered as of this publication date; the horizons already past (90, 60, 45 days) are scored only from probabilities that the record recorded before any outcome was known and before this plan was published. The market side of that record is fingerprinted and timestamped daily at scopes.com/integrity, and this page is archived to the Wayback Machine on publication.

The specification

Universe
Every race on the polls roster that carries both a liquidity-gated market and a published Scopes Fair Value under the standing eligibility rule. The Senate-control and House-control markets are included as their own rows. A race that never cleared the composite gate (≥2 models, or ≥3 qualifying pollsters for the Poll Average) at a horizon is excluded at that horizon, not filled.

Horizons
90, 60, 45, 30, 21, 14, 7, 3 and 1 days before Election Day (November 3, 2026).

Probability used (pre-registered rule)
For each source at each horizon, the as-published probability from the last snapshot strictly before the horizon time: the market; the Scopes Fair Value; each named model where present; and the Scopes Poll Average where it was live. A composite is never recomputed with later methodology; the value the record holds is used, with its methodology_version. Named-model lines are limited to rights-clean, independent sources (open or display licensed, not a market-input model, and in the current composite slate), consistent with the standing source-rights policy; the headline market versus Scopes Fair Value comparison is unaffected.

Score
Brier per race per horizon per source against the resolved outcome (0/1); the mean across races per horizon; and the paired market-minus-composite difference with a cluster bootstrap by race (1,000 resamples) for the confidence band.

Sample gate
8 races with both sides available at a horizon. Below 8, the horizon's point is plotted with an early-data label and reads "insufficient power, no verdict." This gate was set from the availability audit below (roster of ~13 races; ~10 clear both sides at the mature horizons), not chosen after seeing results.

Output
One chart (x = days before Election Day, y = mean Brier, one line per source, shaded band on the paired difference) plus the table, and an expandable per-race table so any single race's curve is inspectable. Rendered server-side with the numbers in the HTML, in the dataviz house palette, light and dark.

Verdict templates (filled mechanically after resolution)

Chosen only by the sign and the significance of the paired difference at that horizon. The braces are filled from the computed numbers; no other text is added.

60 days out

At 60 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).

30 days out

At 30 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).

7 days out

At 7 days, the market’s mean Brier was {lower / higher / no different} than the {composite} by {delta} ({significant / not significant} at 95%, cluster bootstrap by race, n = {races}).

Data availability at registration

The collectors began on July 15, 2026 and the roster grew over the summer, so the earliest horizon is thin. We disclose that here rather than backfilling it. Races with both a market snapshot and a composite value within 24 hours before each past horizon:

HorizonDateRaces with both sidesStatus
90dAug 51Below the gate: no verdict
60dSep 410Meets the gate
45dSep 1910Meets the gate
30dOct 4—Future at registration; fills at 6h cadence
21dOct 13—Future at registration; fills at 6h cadence
14dOct 20—Future at registration; fills at 6h cadence
7dOct 27—Future at registration; fills at 6h cadence
3dOct 31—Future at registration; fills at 6h cadence
1dNov 2—Future at registration; fills at 6h cadence

Exclusions, pre-registered: governor races (no published composite: no forecast-model coverage and the Poll Average is below its pollster gate) and any Kalshi-only race whose composite is gated are excluded until and unless they clear the gate. Snapshot cadence for the market and the composite through Election Day is at least every 6 hours (ingest at 00, 06, 12, 18 and 07 UTC, with the composite recomputed each run).

The selection rule

Both sides come from the append-only record. The market side is price_snapshots (per contract, keyed to the race outcome); the composite and each model are benchmark_probs rows keyed by source (scopes-composite, scopes-poll-avg, statsheet-model). At each horizon each source's value is the last row strictly before the horizon time. The canonical availability query is pinned below; the scoring code and its methodology version are published with the results.

Horizon selection (per source): last snapshot strictly before the horizon

select implied_prob, captured_at
from <price_snapshots | benchmark_probs>
where <contract_id | outcome_id> = :id
  and captured_at < :horizon_time
order by captured_at desc
limit 1;

What this can't test

One cycle is a small sample. Where a horizon lacks the races for a stable estimate, its verdict reads "insufficient power, no verdict," by rule, rather than overclaiming on thin data. Brier measures calibration and resolution against a binary outcome; it does not adjudicate why a source was right or wrong. Those questions are the separate, already-registered claims above.

Follow the inputs as they fill in on the poll desk and the scorecard. Results here are computed after November 3, 2026.