Methodology & changelog
How the Scopes Fair Value is computed
Each desk's Scopes Fair Value comes from a benchmark that is independent of the betting-market price; that independence is what makes a divergence meaningful. Every /api/v1/series response carries a methodology_version; the table below is the key for those strings.
| Desk | Method | methodology_version |
|---|---|---|
| Rates | Futures-implied probabilities (FedWatch method) and macro nowcasts (CPI, GDP).CME futures / nowcast models, not the betting-market price. | benchmark source per event (e.g. futures-implied, nowcast) |
| Polls | Scopes Poll Composite of independent election-forecast models.Third-party forecast models. | composite-v1 |
| Sports | De-vigged sharp sportsbook line (Pinnacle), margin removed.A different, sharp market, not Kalshi. | sharp-book-devig |
| Stocks | Realized-volatility statistical model: P(range) from the index return distribution.Historical realized vol, NOT the options surface the contract is priced off. | realized-vol-statistical (was cboe-options-implied before 2026-08-13) |
| Energy | Realized-volatility statistical model on WTI: P(range) from the WTI return distribution.WTI realized vol, not the ETF-options surface. | realized-vol-statistical (was cboe-options-implied before 2026-08-13) |
| Weather | NWS-forecast model: P(high in band) from the daily-high forecast under a normal with a lead-time error σ.NWS forecast (api.weather.gov), not the venue’s settlement source (The Weather Company). | nws-forecast |
| Catastrophe | Negative-binomial count model: P(season count > N), mean = curated NOAA/CSU seasonal outlook for that basin/metric, dispersion calibrated to basin climatology, conditioned on the season-to-date count. v1 = East/Central Pacific named-storm & hurricane counts + Atlantic named storms (Atlantic hurricane/major-count & landfall markets to follow).NOAA/CSU outlooks + basin climatology (HURDAT2), not the venue’s settlement source (the NHC post-season count). | nb-noaa-climatology-v1 |
Versioning rule
A change to how a fair value is computed bumps the methodology_version: we do not silently rename or redefine a method. Series points keep the version under which they were computed, so a pinned point-in-time (as_of) query remains reproducible and interpretable for good. Data errors are handled as restatements, dated and noted (see the terms).
Scopes Poll Average: v1
A rules-based benchmark for election races: objective inclusion filters select qualifying public polls, one-per-pollster dedupe removes the house-effect correlation, the median of two-party margins is the aggregate, and a versioned sigma schedule translates that margin into a win probability. No per-poll human judgment enters the number. This is a measurement benchmark (not investable).
The v1 rules (verbatim):
- Measurand: the two-party margin (Rep − Dem) in the named race; then P(win) for each party.
- Inclusion: a public release that discloses pollster, field dates, sample size, population, and mode; polls the actual named matchup; is not campaign- or party-sponsored; and has n ≥ 400. Window: the trailing 21 days, or the latest 5 qualifying polls, whichever is larger.
- Dedupe: one poll per pollster (most recent), so a prolific pollster cannot dominate the median.
- Multiple matchups (v1 refinement, 2026-08-23): when one pollster tests more than one Democratic matchup for the same seat (the nominee may be unsettled), the poll's margin is the mean of those matchups' Rep−Dem margins, one deterministic value per pollster, mapping a candidate poll onto a party-control market with no per-poll judgment.
- Aggregation: the median of the deduplicated margins.
- Publication gate: at least 3 distinct pollsters, or the race shows its polls but no average and no fair value.
- Margin → probability: P(Rep) = Φ(median margin ÷ σ), with the horizon σ schedule 5.5 pts at 60+ days to the election, 4.5 at 15–59 days, 3.5 under 15 days (basis: documented historical state-poll error). Undecideds are left unallocated; an exact tie is 50%.
SAMURAI. Specified in advance (rules published here before use); Appropriate (a poll average is the standard read of a horse-race margin); Measurable + Unambiguous (a deterministic median + formula: same inputs always yield the same number); Reflective (qualifying public polls of the named race); Accountable/Owned (Scopes publishes and owns the methodology, versioned). It is a measurement benchmark, not investable.
Versioning. These v1 rules are immutable; they never change retroactively. Any improvement launches as v2 alongside v1, with its own series; each computed value carries its methodology_version(poll-avg-v1). Every input poll is stored with its original publisher release link, so any figure is reproducible.
Race eligibility (which races are on the desk, a rule, not a choice):
The roster is registry logic, never editorial. A race joins the polls desk when both hold: (a) a liquidity-gated market exists on at least one tracked venue, and (b) the Scopes Poll Average publication gate clears: ≥ 3 distinct qualifying pollsters in the window. If (b) later lapses, the race auto-suspends: the board shows “insufficient polling” instead of a fair value, rather than being quietly dropped. The same rule governs Senate, Governor, and House races alike; House districts are admitted only when they clear it, never by hand.
Recession probability (NY-Fed yield curve)
The recession desk's Scopes Fair Value is the probability of a US recession within the next twelve months, computed from the Treasury yield curve using the probit model the Federal Reserve Bank of New York publishes monthly (Estrella & Mishkin). It is a measurement benchmark (not investable), and it is a reference: it is shown for context and is not scored on the public record (see the note below).
The v1 rules (verbatim):
- Input. The month-average spread of the 10-year Treasury constant-maturity yield minus the 3-month Treasury bill (secondary market), in percentage points, the public FRED series
T10Y3M, averaged over the latest calendar month. - Formula.
P(recession) = Φ(α + β · spread), where Φ is the standard normal CDF and the coefficients are the NY Fed's published probit estimatesα = −0.5333,β = −0.6330(estimated January 1959–December 2009). Recomputing from the current spread reproduces the NY Fed's published figure (e.g. a 0.7822 spread → 15.19%). - Provenance. Coefficients and method from the NY Fed's “The Yield Curve as a Leading Indicator”; the spread from FRED (original publishers). No third-party model is ingested.
Reference, not scored. The model estimates a rolling twelve-month probability, while the tracked contracts resolve on fixed dates (“recession by end of 2026 / 2027”), so the horizons do not line up. Publishing a market-vs-benchmark divergence as a scored result would be unfair, so recession is shown as context and is never recorded on the scorecard. A horizon-matched treatment, if built, would launch as v2.
SAMURAI. Specified in advance (a published probit), Appropriate (a recession estimate for recession markets), Measurable and Unambiguous (one formula, one public input), Reflective, Accountable/Owned (the NY Fed's method, computed and versioned by Scopes). It is a measurement benchmark, not investable.
Versioning. These v1 rules are immutable; they never change retroactively. Any improvement launches as v2 alongside v1; each computed value carries its methodology_version (recession-nyfed-v1).
Changelog
2026-08-26
Recession benchmark (NY-Fed yield curve) v1 added
Added a recession desk: P(US recession within 12 months) from the NY Fed’s published probit on the 10yr−3mo Treasury spread (FRED T10Y3M), methodology_version recession-nyfed-v1 (α=−0.5333, β=−0.6330). Recompute matches the NY Fed’s own published figure to 4dp. Mapped the recession-by-date contracts on both venues (Kalshi KXRECSSNBER, Polymarket). It is a measurement benchmark (not investable) and, for v1, a REFERENCE that is not scored on the record: the model’s rolling 12-month horizon does not match the contracts’ fixed resolution dates, so a scored divergence would be unfair. A horizon-matched treatment, if built, launches as v2. See the Recession probability section above for the verbatim rules and SAMURAI statement.
2026-08-23
Scopes Poll Average: multi-matchup rule (v1 refinement)
Defined how the poll average handles a pollster that tests more than one Democratic matchup for the same seat (common before a primary settles the nominee, while the market is party-control): the poll’s margin is the mean of the tested matchups’ Rep−Dem margins, one deterministic value per pollster, no per-poll judgment. This is a same-day refinement of previously-underspecified behavior; it changes no already-computed value (every entered poll to date was single-matchup), so the series stays methodology_version poll-avg-v1. Both matchup margins are stored for reproducibility.
2026-08-22
Scopes Poll Average v1 added (Michigan Senate 2026 pilot)
A new rules-based poll-average benchmark (methodology_version poll-avg-v1): objective inclusion filters select qualifying public polls, one-per-pollster dedupe removes house-effect correlation, the median of two-party margins is the aggregate, and a versioned sigma schedule (5.5/4.5/3.5 pts by horizon) translates margin to P(win). At least 3 distinct pollsters or no average is published. It joins the Scopes Poll Composite as a named model. Piloted on Michigan Senate 2026; populated from public polls entered with their original-publisher release links. See the Scopes Poll Average section above for the verbatim rules and SAMURAI statement.
2026-08-18
Weather Scope: per-city forecast-error σ recalibration (pre-launch)
The weather fair value now uses a PER-CITY forecast-error σ (σ = base + growth·lead) instead of one global curve. The global σ was far too wide for low-variance coastal cities (a liquid Los Angeles market implied σ≈1°F at ~1.5 days while the global model used ~2.9°F), which smeared probability across temperature bands and manufactured divergences. Coastal/marine-moderated cities (LA, Miami) get a tight σ; continental cities (Chicago, NY) a wider one. Weather remains not-live (dormant) while we also resolve an apparent forecast-center offset (NWS gridpoint vs. the venue’s settlement station) and let flags resolve — the desk flips live only once the fair tracks liquid markets and the scorecard shows no systematic bias.
2026-08-17
Catastrophe Scope: live-storm landfall layer (GEFS ensemble)
Added the live-storm landfall fair value. When an Atlantic storm is active, the Scopes fair value for a per-city strike is the GEFS ENSEMBLE strike fraction — the share of the 31 member tracks (AP01–AP30 + control, from NHC’s public a-deck) passing within the strike radius (~65 nm) at hurricane intensity (≥64 kt) — blended toward the remaining-season climatological prior by how in-play the storm is for that city. This is raw model guidance, deliberately NOT NHC’s official forecast (the cone) and NOT the realized-landfall settlement (NHC best-track), so it can diverge from both the market and NHC. Kalshi lists per-city landfall markets only when a storm threatens; the storm room (/catastrophe/storm) and /api/v1/landfall show the fair value whenever a storm is active, with the market price alongside once a market is live. The ensemble→probability mapping is a v1 heuristic to calibrate against resolved storms; version gefs-ensemble-v1.
2026-08-17
Catastrophe Scope: Atlantic per-city landfall climatology foundation
Added the independent fair-value engine for Atlantic per-city hurricane landfall (the insurance/cat-bond-aligned view). Landfalls near a city are modeled as a Poisson process at the climatological annual rate λ = 1 ÷ return period (return periods from HURDAT2 track history), so P(strike this season) = 1 − e^(−λ); a rest-of-season figure scales λ by the share of Atlantic landfall risk still ahead (risk clusters Aug–Oct). Kalshi lists per-city landfall markets only when a storm threatens (they are storm-triggered, not always-on), so this ships first as a standalone climatology reference (/catastrophe/landfall); the market-vs-fair comparison layer snaps onto this engine when live markets appear. Independent of the venue’s settlement source (NHC best-track). Return periods are v1 climatology placeholders to refine with a live HURDAT2 track parse.
2026-08-16
Catastrophe Scope added: Pacific seasonal storm counts vs. a negative-binomial model
The catastrophe desk (v1: East & Central Pacific seasonal named-storm and hurricane count markets — Kalshi does not currently offer Atlantic seasonal-count markets; an Atlantic landfall desk is planned) computes its Scopes Fair Value from a negative-binomial count model. The contracts are “more than N” (strictly greater); the model’s mean is the curated NOAA CPC seasonal outlook for that basin and metric, and the dispersion (NB size) is calibrated to the basin’s historical variance, so the tails match how variable those seasons actually are — a Poisson pinned to the mean would underweight the tails the market prices (Pacific counts are ENSO-driven and over-dispersed, especially Central Pacific). The distribution is conditioned on the season-to-date count (NHC): mid-season it prices the remaining season. This is deliberately a different source than the venue’s settlement (the NHC post-season count), so the fair value is independent. The forecast mean and season-to-date count are curated inputs refreshed a few times a season; the dispersion is a v1 climatology calibration to refine as resolved seasons accumulate. Collecting dormant first — the desk shows as “coming” until the outlook values and calibration are validated.
2026-08-16
Weather Scope added: daily high-temperature vs. NWS forecast
The weather desk (v1: daily city high-temperature markets) computes its Scopes Fair Value from the NWS forecast (api.weather.gov, public domain) as P(high in a band) under a normal centered on the forecast high with a lead-time error σ (~2°F same-day rising to ~6°F at a week). This is deliberately a different source than the venue’s settlement (The Weather Company), so the fair value is independent. Near-settled same-day contracts are held out of the scorecard and board — the market has collapsed to the nearly-known high while the forecast σ has not, which would produce spurious gaps. The σ is a v1 heuristic and will be tuned empirically as resolved outcomes accumulate.
2026-08-13
Stocks & energy fair value → independent realized-vol model
The stock and energy desks moved their headline Scopes Fair Value from an options-implied value (cboe-options-implied) to an independent realized-volatility statistical model (realized-vol-statistical). The options-implied fair was derived from the same options surface the Kalshi contract is priced off, so market and fair value agreed almost by construction. The statistical model estimates the return distribution from realized volatility — a source the traded price does not echo — so the divergence now reflects the market pricing a wider distribution than realized vol (the variance risk premium). The options-implied value is retained as a reference column. Series points captured before this date remain labeled cboe-options-implied, so pinned as_of queries stay accurate.
Machine surfaces: the OpenAPI spec and llms.txt. Every series point is a calculation with a methodology_version. Scopes is research and information only, not investment advice.