Scopes

Research · pre-registered study

Five claims about prediction markets and polls — and how the record will test them

Prediction markets are quoted as if they settle the question. Pollsters push back: the markets are downstream of polls, they inherit poll bias, and they reflect who is doing the betting. Those are testable claims. We state five of them here — before the 2026 midterms resolve — with the exact test, threshold, and verdict line for each, and the record fills in the numbers mechanically after the races settle. No prediction of which claim will hold; the verdict templates are fixed in advance.

The marquee question is the pollster's own: whether the market is mostly the polls, or something more. The honest, computable form is how much the market adds beyond the Scopes Fair Value (H1, H5) — not the uncomputable counterfactual of a world with no polls.

Pre-registered September 5, 2026. The claims, methods, thresholds, and verdict templates below are fixed as of this date and archived at publication. Results are computed mechanically from the tracked record after each race resolves, into the Midterms Scorecard — no discretionary analysis, no re-specification. Independent timestamped snapshot: archive.ph/hj3DW (5 Sep 2026 09:29 UTC).

H1 · Downstream

The claim
Prediction markets are downstream of polls — they move because polls move. (Paraphrasing the pollster Kristen Soltis Anderson, NYT.)

The test
Lead–lag between market-price moves and Scopes Poll Composite moves, per race.

Method (pre-registered)
Cross-correlation of the two change series at lags of 1h, 6h, and 24h, per race, over the tracked window; report the peak-correlation lag (sign = which series leads) and the pooled median lead across races. Tests lead–lag against the Scopes Composite specifically — not against all polling.

Verdict line (filled after resolution)
H1 — Downstream: the market {led / lagged / moved concurrently with} the Composite by a pooled median of {X}h (peak cross-correlation at lag {L}, n = {races}).

Result: pending resolution.

H2 · Garbage-in, garbage-out

The claim
If the polls behind the bets are systematically biased, the markets inherit that bias — garbage in, garbage out.

The test
Correlation of market resolution error with Composite resolution error across races.

Method (pre-registered)
Per race, error = |final probability − realized outcome (0/1)| for the market and for the Composite. Report the cross-race correlation of the two error series (Pearson + rank), and the share of races where both erred in the same direction.

Verdict line (filled after resolution)
H2 — Garbage-in: market and Composite errors correlated at r = {r} ({direction}); they erred together in {pct}% of races (n = {races}).

Result: pending resolution.

H3 · Team bias

The claim
Bettors bet the team they want to win, so markets reflect the participants’ biases.

The test
Signed divergence (market − Composite) by party, across all races — mean and significance.

Method (pre-registered)
Per race, signed gap = market − Composite for the same party outcome; pool by party; report the mean signed gap, a two-sided significance test, and cluster-bootstrap CIs by event family. NOTE: a lean is interpretable only with H2/H4 — a party lean that proved right at resolution is information, not bias; a lean that proved wrong is bias.

Verdict line (filled after resolution)
H3 — Team bias: the market leaned {party} by a mean {pts} pts vs the Composite ({significant / not significant}); paired with H4, that lean was {right / wrong} at resolution (n = {races}).

Result: pending resolution.

H4 · Who was right on divergence

The claim
When the market and the polls disagreed, which side proved right?

The test
Polls-desk resolution split and paired Brier score, on flagged divergences.

Method (pre-registered)
On divergences flagged above threshold, the mechanical scorecard records which side resolved correct; report the split and the paired Brier score (market vs Composite), cluster-bootstrapped by event family for CIs.

Verdict line (filled after resolution)
H4 — Who was right: on {flags} flagged divergences, {market / the Composite} was right {pct}% of the time; paired Brier {b_market} vs {b_composite} (n = {races}).

Result: pending resolution.

H5 · Added information

The claim
Does the market carry information beyond the polls — the answerable form of "what would the markets do without polls?"

The test
Does the market beat the Composite on Brier when the two disagree beyond the flag threshold?

Method (pre-registered)
Restrict to snapshots where |market − Composite| exceeds the flag threshold; compute each side’s Brier at resolution over that subset; report the difference and a cluster-bootstrap CI. This is the defensible form of the counterfactual: not "markets without polls," but "how much the market adds beyond the Composite."

Verdict line (filled after resolution)
H5 — Added information: where they disagreed, the market’s Brier was {lower / higher / no different} than the Composite’s by {delta} ({significant / not significant}, n = {races}).

Result: pending resolution.

What this can't test

The pollster's strongest point is one we concede: you can see a poll's author, methodology, and incentives, but not who is on each side of a trade or why. Price data measures the market's outputs — information, bias, accuracy — never the transparency of its participants. And this is a single cycle: where a hypothesis lacks the sample size for a stable estimate, its verdict line reads "insufficient power — no verdict," by rule, rather than overclaiming on thin data.

Follow the numbers as they fill in on the scorecard and the poll desk.