Research · pre-registered study
Are prediction markets calibrated on CPI? Pre-registered
The Scopes Desk
A common claim is that prediction markets are well-calibrated on scheduled macro releases. This plan pre-registers one mechanical test of it: do the Kalshi CPI contract probabilities match how often those outcomes actually occur, and does that calibration tighten as the print nears? We score the full strike ladder against the realized print at five fixed horizons before each release, with the method, the universe, the sample gate, and the verdict lines all fixed here before the prints resolve.
This is a rolling panel: one CPI release is a single draw (every strike in a print resolves off one number), so a verdict needs many prints. It accrues monthly and is not expected to reach its sample gate until 2027. The headline here is the registration itself, timestamped before any tested print, not a near-term result.
The specification
Universe
The Kalshi CPI contract ladders: KXCPI (month-over-month) and KXCPIYOY (year-over-year). Every strike in a release that clears the liquidity gate is scored; a strike that never clears the gate at a horizon is excluded at that horizon, not filled. The realized outcome for each strike is set from the Bureau of Labor Statistics print, read from FRED (CPIAUCSL for month-over-month, CPIAUCNS for year-over-year) by the existing release resolver.
Horizons
21d, 14d, 7d, 3d, 1d before each CPI release (measured back from the scheduled release time). Thirty days out the contracts are too thin to score, so the window starts at 21 days.
Probability used (pre-registered rule)
For each strike at each horizon, the as-published market probability from the last snapshot strictly before the horizon time. Nothing is recomputed or interpolated; the value the record holds is used.
Score
Two measures, pooled across the strike ladder and across prints: the mean Brier of the contract probability against the 0/1 outcome, and the expected calibration error (ECE) over 5 equal-width reliability bins. Confidence bands are a cluster bootstrap by print (1,000 resamples), because one print determines every strike in it.
Sample gate
8 resolved prints with scored strikes at a horizon. Below 8, that horizon reads "insufficient power, no verdict." The gate is in prints, not strikes, for the same reason: the strikes within a print are not independent.
Output
A reliability table (forecast probability vs. observed frequency per bin) and a mean Brier and ECE per horizon, each with its cluster-bootstrap band, rendered server-side with the numbers in the HTML.
Verdict templates (filled mechanically after resolution)
Chosen only by where the calibration-error confidence band falls relative to the pre-registered band (ECE at or below 0.05 reads well calibrated; a band entirely above it reads miscalibrated; a band straddling it reads inconclusive). The braces are filled from the computed numbers; no other text is added.
14 days out
At 14 days, across {prints} prints and {strikes} contract-strikes, the market’s mean Brier was {brier} and its calibration error was {ece} ({well calibrated / miscalibrated / inconclusive} at 95%, cluster bootstrap by print; well-calibrated band is ECE at or below 0.05).
3 days out
At 3 days, across {prints} prints and {strikes} contract-strikes, the market’s mean Brier was {brier} and its calibration error was {ece} ({well calibrated / miscalibrated / inconclusive} at 95%, cluster bootstrap by print; well-calibrated band is ECE at or below 0.05).
Data availability at registration
CPI contract capture was re-enabled in October 2026. The panel begins at the first release for which the horizon snapshots exist; no earlier horizon is backfilled. The first release with full horizon coverage (snapshots at all five horizons recorded before any outcome was known) is the first November 2026 CPI release and later. Earlier releases, if any, contribute only the horizons that were actually captured before the print, and are labeled as partial. Snapshot cadence for the CPI contracts is at least every six hours.
The selection rule
Both the market series and the outcome come from the append-only record. The market side is price_snapshots (per CPI contract, keyed to its strike); the outcome is set by the release resolver from FRED. At each horizon each strike's value is the last row strictly before the horizon time. The scoring code and its methodology version are published with the results.
Horizon selection (per strike): last snapshot strictly before the horizon
select implied_prob, captured_at from price_snapshots where contract_id = :strike_contract_id and captured_at < :horizon_time order by captured_at desc limit 1;
What this can't test
- It is slow. A verdict needs 8 prints, so the earliest readout is 2027. Until then every horizon reads "insufficient power, no verdict" by rule.
- Brier and calibration error measure whether the probabilities were right, not why a price was where it was.
- It does not yet compare the market to the Cleveland Fed inflation nowcast. That benchmark's terms authorize reproduction for noncommercial, personal, or educational use only, so the head-to-head is deferred until that use is licensed or run under a noncommercial program. The comparison, if added, will be a logged amendment registered before any nowcast value is used.
Follow the CPI contracts as they fill in on the rates desk. Results here are computed after each print resolves.