Scopes

The independent record

Calibration.

How accurate are the markets, really? Every number here is computed from resolved flags: the mechanical record of when a market price and the Scopes Fair Value diverged, tracked to outcome. No opinions, no projections; the statistics decide what can and can't yet be said.

612
Resolved flags
364
Market right
248
Scopes Fair Value right

Is the difference real yet?

248/612 resolved to the Scopes Fair Value side. If both sides were equally good, a split this uneven or more would occur by chance < 1% of the time (p = < 0.01).

The difference is statistically significant at the 95% level.

At the current flag rate, a true 60/40 skill difference becomes detectable around n ≈ 194 resolved flags; a 55/45 difference around n ≈ 783.

Brier comparison

0.311
Market: mean Brier
n = 612
0.358
Scopes Fair Value: mean Brier
n = 612

Mean Brier difference (market minus fair value): -0.047 [95% CI -0.063, -0.032]

The interval excludes zero: the market has shown the more accurate probabilities so far.

The Brier score is the mean squared error between a probability and the 0/1 outcome, lower is more accurate. Probabilities are taken as recorded at flag time, not revised after the fact. The interval is a cluster bootstrap: correlated outcomes, e.g. multiple flags on the same election cycle, are resampled together so they cannot manufacture false precision. This is a paired forecast-comparison in the spirit of Diebold–Mariano. What the scorecard’s split does and doesn’t mean →

By desk

DesknRight (fair / market)Binomial pBrier: marketBrier: fair
Rate Scope2— · n < 10
stocks4522 / 231.000.3650.447
energy404154 / 250< 0.010.3180.304
weather16170 / 910.110.2760.469

Explore the record

Filter the resolved record by desk, event family, minimum Scopes Divergence, and date range, every slice recomputes the same binomial test and paired Brier comparison. Slices below 10 flags report counts only, no verdict.

Loading the record…

Flags are deduplicated (one per binary event) and grouped into event families for testing; families, not individual flags, are the unit of statistical resampling. Sports flags are the closest to independent observations; election flags within one cycle are treated as correlated.