Research · pre-registered study
Frontier models vs the market: 2026 Senate control
A pre-registered contest. Before Election Day, every frontier model in our roster forecasts the winner of each contested 2026 U.S. Senate race, plus the one question that decides everything: who controls the Senate. Each model gets our own Scopes Poll Composite as context. The forecasts are locked before the results are in, then scored by a proper scoring rule against the outcomes and against the market.
Scopes is the scorekeeper, not a forecaster. The models forecast; the market forecasts; we run the contest, lock the predictions, and keep the append-only record. The prompt and the exact context shown to each model are published below so any score here can be reproduced.
What we registered in advance
- Across the contested races, does any model's mean Brier score beat the market's (Kalshi per-race implied)?
- Pick record: of the races, how many did each model call correctly vs the market?
- On the single question that matters most, Senate control, does any model beat the Kalshi control market?
- Does our poll composite let a model match a market that already prices the polls?
Registered expectation: with roughly eight contested seats plus one control question, this is underpowered, so we treat any separation as directional, not conclusive. Control is a single outcome the whole field forecasts. The market prices the polls efficiently; a model failing to beat it is a legitimate, reportable result.
The contest
Pre-registration opens at the field lock (before Election Day). Once the forecasts are locked they appear here immediately, and scores fill in as the races are called.
The prompt, put to every model
The identical prompt, with each race's context filled in. The model is sampled five times; the mean probability is the locked forecast. It never sees the market price: the market is the benchmark, not an input. The seats use the per-race prompt; the control question uses the control prompt.
You are forecasting a 2026 U.S. Senate race that has not yet been decided. Give your best estimate of the probability that the REPUBLICAN candidate wins the seat.
Race: {race}
Scopes Poll Composite (our poll-based estimate of P(Republican wins)): {composite}
{fundamentals}
Use what you know about this race together with the poll signal above. Answer with a single JSON object and nothing else:
{"p_republican_wins": 0.0-1.0, "confidence": 0.0-1.0, "reason": "<=240 chars"}You are forecasting which party will control the U.S. Senate after the 2026 elections. Give your best estimate of the probability that the REPUBLICANS control the Senate.
Scopes Poll Composite (our poll-based estimate of P(Republicans control)): {composite}
{fundamentals}
Use what you know about the 2026 Senate map together with the poll signal above. Answer with a single JSON object and nothing else:
{"p_republican_wins": 0.0-1.0, "confidence": 0.0-1.0, "reason": "<=240 chars"}Context is our own Scopes Poll Composite (a scopes-derived value), not a third-party aggregator. The market benchmark is the Kalshi (and, where shown, Polymarket) implied probability from our own poll data, captured at lock.