Research · pre-registered study
Frontier models vs the market: the 2026 MLB postseason
A pre-registered contest. Before each round of the 2026 postseason, every frontier model in our roster forecasts the winner of each series, with public context: records, probable starters, rest, and injuries. The forecasts are locked before a pitch is thrown and scored, once the series settle, by a proper scoring rule against the actual outcomes and against the market.
Scopes is the scorekeeper, not a forecaster. The models forecast; the market forecasts; we run the contest, lock the predictions, and keep the append-only record. The prompt and the exact context shown to each model are published below so any score here can be reproduced.
What we registered in advance
- Over the ~11 postseason series, does any model's mean Brier score beat the market's?
- Straight pick record: of the series, how many did each model call correctly vs the market?
- Does context (records, rotation, rest, injuries) let a model match a market that already prices everything?
Registered expectation: at roughly eleven series this is underpowered, so we expect no statistically significant separation. This is a demonstration of the mechanism and a calibrated look, not a conclusive ranking. The market is a hard baseline; a model failing to beat it is a legitimate, reportable result.
The contest
Division Series
| anthropic claude-opus-5 | P(Cleveland Guardians) 58.8% | |
| google gemini-3.8-flash | P(Cleveland Guardians) 56.1% | |
| openai gpt-6-astra | P(Cleveland Guardians) 56.2% |
| google gemini-3.8-flash | P(Los Angeles Dodgers) 54.5% | |
| openai gpt-6-astra | P(Los Angeles Dodgers) 58.0% | |
| xai grok-4.6 | P(Los Angeles Dodgers) 60.2% |
| anthropic claude-opus-5 | P(Tampa Bay Rays) 55.6% | |
| google gemini-3.8-flash | P(Tampa Bay Rays) 56.3% | |
| openai gpt-6-astra | P(Tampa Bay Rays) 58.6% | |
| xai grok-4.6 | P(Tampa Bay Rays) 59.4% |
| anthropic claude-opus-5 | P(Milwaukee Brewers) 60.0% | |
| google gemini-3.8-flash | P(Milwaukee Brewers) 58.6% | |
| openai gpt-6-astra | P(Milwaukee Brewers) 63.2% | |
| xai grok-4.6 | P(Milwaukee Brewers) 63.7% |
The prompt, put to every model
The identical prompt, with each series' context filled in. The model is sampled five times; the mean probability is the locked forecast. It never sees the market price: the market is the benchmark, not an input. Each locked series above carries the exact context that was shown.
You are forecasting a Major League Baseball postseason series that has not yet been played. Give your best estimate of the probability that the FIRST team wins the series.
Series: {team_a} vs {team_b}, best-of-{games}
Home-field advantage: {home_team}
Regular-season records: {team_a} {record_a}, {team_b} {record_b}
Season head-to-head: {h2h}
Probable starters: {team_a}: {rotation_a}; {team_b}: {rotation_b}
Rest / fatigue: {team_a}: {rest_a}; {team_b}: {rest_b}
Injuries (IL): {team_a}: {il_a}; {team_b}: {il_b}
Use what you know about these teams and postseason baseball, and weigh rotation, rest, and injuries (a spent ace or a tired team on no rest matters in a short series). Answer with a single JSON object and nothing else:
{"p_team_a_wins": 0.0-1.0, "confidence": 0.0-1.0, "reason": "<=240 chars"}