Research · the read-off · standing record
Which model reads the event best
For every open event on the scorecard there are two independent estimates of the same probability: the price implied by a prediction market, and an independent fair value. We ask each frontier model the identical question, before the event resolves: which of the two is better calibrated? Once the event settles, the answer is known, and we record which estimate each model chose and whether it matched the one that proved closer. This page is the running tally.
The models are not forecasting the outcome and are not advising a wager. They are grading two estimators. The prompt, the inputs, and the raw outputs are published below so any score here can be reproduced.
The standing record
The standing record scores only the numbers-only reads: the identical prompt every model sees on every desk, with nothing but the two probabilities. It is the one cohort where cross-model accuracy is comparable, because the prompt is held constant. Reads run with an added reference block are a separate experiment, scored below and never mixed in here.
Accuracy is the share of a model's decided reads where the estimate it chose proved better calibrated (closer to the resolved outcome). Reads where the two estimates were equidistant are a Push and are excluded. A model needs at least 20 decided reads before it carries an accuracy figure; below that it shows its counts only.
Models differ sharply in how often they abstain, so a low decided count is often about caution, not skill. On a numbers-only read, given just the two probabilities, some models judge they have no basis to prefer either estimate and answer indeterminate on most events; others reason from what they know about the event and commit. Abstaining is a legitimate answer, not a wrong one: a model that declines when it has no basis is not penalized, but it also accrues few decided reads and may carry no accuracy figure here. Whether committing more often means reading better is one of the questions this record exists to answer.
| Model | Matched better est. | Decided | By choice |
|---|---|---|---|
| anthropic claude-opus-5 | 72.7% (16/22) | 22 | market 16/22 |
| google gemini-3.8-flash | early | 13 | market 9/12 · fair 0/1 |
| xai grok-4.6 | early | 1 | market 0/1 |
| openai gpt-6-astra | early | 0 | — |
| deepseek deepseek-v4-pro-0813 | early | 0 | — |
Does context make a model commit?
A separate experiment, not part of the standing record above. On the desks whose fair-value inputs are rights-clean and public domain (weather, catastrophe, rates), the prompt also carries a short reference block: the forecast, climatology, or macro figure the fair value was built from, published verbatim with each read. This panel compares each model's numbers-only reads to its enriched reads on those desks. The first thing context moves is the commit rate: how often a model answers at all instead of abstaining.
| Model | Commit rate: numbers-only → enriched | Accuracy (n.o. / enriched) |
|---|---|---|
| xai grok-4.6 | 9.6% (8/83) → 100.0% (33/33) | 0.0% (0/1) / 57.1% (12/21) |
| google gemini-3.8-flash | 77.1% (27/35) → 100.0% (25/25) | 69.2% (9/13) / 68.8% (11/16) |
| anthropic claude-opus-5 | 74.5% (35/47) → 100.0% (30/30) | 72.7% (16/22) / 75.0% (15/20) |
| deepseek deepseek-v4-pro-0813 | 66.7% (2/3) → 100.0% (9/9) | — / 100.0% (1/1) |
| openai gpt-6-astra | 4.8% (4/83) → 75.7% (28/37) | — / 58.8% (10/17) |
Read commit rate first. Accuracy here is on small samples and is not a rating; it is shown only to check that committing more is not the same as reading worse. The enriched desks are the public-domain ones, not the numeric desks (energy, stocks) where the most cautious models abstain most, so this does not touch those models' standing figures.
The question, put to every model
The identical prompt is sent to every model, with the event, the outcome, and the two probabilities filled into the placeholders. Each model is sampled 5 times at temperature 0.3; the majority read is recorded. A model may answer indeterminate when the two numbers give it no basis to prefer either; those abstentions are recorded for transparency but never scored.
The {reference} line is empty on numbers-only reads. On the public-domain desks it carries the reference block described above, and its exact text is published with each enriched read in the list below.
You are grading two independent probability estimates for a FUTURE event that has not yet occurred. You are NOT predicting the outcome and NOT advising a bet.
Event: {event}
Outcome being judged: {outcome}
Two estimates of the probability of this outcome are on record:
- MARKET (prediction-market implied): {market_pct}
- FAIR VALUE (an independent benchmark: {context}): {fair_pct}
{reference}
Question: which estimate is better calibrated, that is, closer to the true probability of the outcome, the MARKET or the FAIR VALUE? Use whatever you know about this event to judge. If you genuinely have no basis to prefer either, answer "indeterminate" rather than guessing.
Answer with a single JSON object and nothing else:
{"read": "market" | "fair" | "indeterminate", "confidence": 0.0-1.0, "reason": "<=200 chars"}How a read is scored
- A read is recorded before the event resolves and locked with the two probabilities it was shown.
- When the event settles, the outcome is 1 or 0. The estimate closer to that outcome is the better-calibrated one. A read matched the better estimate if the model chose that estimate.
- If the market and the fair value were exactly equidistant, the read is a Push and does not count toward accuracy. An indeterminate answer is an abstention: recorded, but never scored.
- The record is append-only and version-stamped: a model's reads are tallied under the exact version that produced them.
Recent reads, with raw outputs
The most recent reads on record, each with the raw samples the model returned. Expand any row to reproduce the scoring by hand.
Los Angeles 89–90° · openai abstainedOct 1, 4:03 PM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "89-90°F",
"nws_forecast_high_f": 87
}[
{
"read": "indeterminate",
"reason": "The band is 2–3°F above forecast, but plausible forecast-error spreads favor different estimates. Without a lead time or station-specific uncertainty, neither is clearly better calibrated.",
"confidence": 0.8
}
]Los Angeles 87–88° · xai read Market · openOct 1, 4:03 PM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "87-88°F",
"nws_forecast_high_f": 87
}[
{
"read": "market",
"reason": "NWS 87F point forecast has typical 1-2 day RMSE ~2-3F; P(integer high in {87,88}) is then ~25-40%, far closer to 34.5% than 70%.",
"confidence": 0.75
},
{
"read": "market",
"reason": "NWS MAE typically 2-3F even at short lead; 2° bin around a 87° point forecast implies ~25-40% prob, near 34.5% not 70%.",
"confidence": 0.75
}
]Los Angeles 87–88° · openai read Market · openOct 1, 4:02 PM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "87-88°F",
"nws_forecast_high_f": 87
}[
{
"read": "market",
"reason": "An 87°F forecast favors this band, but typical temperature forecast uncertainty makes 70% in a 2°F band too concentrated. Missing lead time and station details limit confidence.",
"confidence": 0.75
},
{
"read": "market",
"reason": "The 87°F forecast supports the band, but typical forecast uncertainty makes 70% in a 2°F interval look overconfident. 34.5% is more plausible; lead time and station are unspecified.",
"confidence": 0.75
},
{
"read": "market",
"reason": "An 87°F forecast supports the band, but typical daily-high uncertainty makes 70% unusually high for a 2°F interval. 34.5% is more plausible, though lead time is unspecified.",
"confidence": 0.75
},
{
"read": "market",
"reason": "An 87°F forecast favors this band, but typical daily-high uncertainty spreads probability beyond two degrees. 34.5% is more plausible than 70%, though lead time and station are unspecified.",
"confidence": 0.75
},
{
"read": "market",
"reason": "Typical daily-high forecast errors spread probability beyond a two-degree band, making 34.5% more plausible than 70%. Unspecified lead time and station limit confidence.",
"confidence": 0.75
}
]SPX 7525–7549.9999 · xai abstainedOct 1, 4:02 PM EDT
[
{
"read": "indeterminate",
"reason": "No data on current SPX, horizon or vol; cannot assess if 25-pt bin density matches 5.5% or 0.1%.",
"confidence": 0.7
}
]SPX 7525–7549.9999 · openai abstainedOct 1, 4:01 PM EDT
[
{
"read": "indeterminate",
"reason": "No settlement date, reference SPX level, or forecast methodology is provided. The two estimates alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, or benchmark methodology is provided. The two quoted probabilities alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, current SPX level, or volatility assumptions are provided. The two estimates alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, reference SPX level, volatility assumptions, or benchmark validation is provided. The two estimates alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, reference SPX level, or benchmark methodology is given, so there is no defensible basis to prefer 5.5% over 0.1%, or vice versa.",
"confidence": 0.99
}
]Los Angeles 82–83° · openai abstainedOct 1, 10:03 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "82-83°F",
"nws_forecast_high_f": 80
}[
{
"read": "indeterminate",
"reason": "An 80°F point forecast alone is insufficient: lead time, station, and forecast-error spread are unspecified. Plausible uncertainty levels could favor either estimate.",
"confidence": 0.9
},
{
"read": "indeterminate",
"reason": "The 80°F forecast alone is insufficient: the probability of 82–83°F depends strongly on forecast error, and no lead time or station-specific uncertainty is provided.",
"confidence": 0.9
},
{
"read": "indeterminate",
"reason": "The 80°F point forecast alone is insufficient: lead time, verifying station, and forecast-error distribution are unspecified and could favor either estimate.",
"confidence": 0.88
},
{
"read": "indeterminate",
"reason": "The 80°F forecast puts the band 2–3°F higher, but lead time, settlement station, and forecast-error distribution are unspecified; these could change which estimate is closer.",
"confidence": 0.88
}
]Los Angeles 82–83° · anthropic read Market · openOct 1, 10:03 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "82-83°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "With NWS high 80°F and ~3°F error, P(82-83) ≈ 15-20%; market's 28.5% is somewhat high but far closer than fair value's implausibly low 4%.",
"confidence": 0.78
},
{
"read": "market",
"reason": "With NWS high 80°F and ~3°F error SD, P(82-83) ≈ 15-20%, far nearer market's 28.5% than fair value's implausibly low 4% for a bin only 2° above forecast.",
"confidence": 0.82
}
]SPX 7625–7649.9999 · xai abstainedOct 1, 10:03 AM EDT
[
{
"read": "indeterminate",
"reason": "No current SPX level, tenor, or vol given for this 25-pt bin; cannot assess which of 9.5% vs 19.1% is better calibrated.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No current SPX level, horizon, vol, or model details given; cannot judge which of 9.5% vs 19.1% is closer to true prob.",
"confidence": 0.55
},
{
"read": "indeterminate",
"reason": "No current SPX level, horizon, or vol given; cannot judge which of 9.5% vs 19.1% is closer to true bin probability.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No data on event horizon, current SPX spot, vol, or model assumptions; cannot judge calibration of either estimate.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No horizon, current SPX level, or vol given; cannot assess which of two unanchored probs is closer to truth.",
"confidence": 0.85
}
]SPX 7625–7649.9999 · openai abstainedOct 1, 10:03 AM EDT
[
{
"read": "indeterminate",
"reason": "No settlement date, market context, or benchmark methodology is provided. The two probabilities alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No event date, benchmark methodology, or calibration history is provided. The two probability estimates alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, forecasting methodology, or calibration history is provided. The two quoted probabilities alone give no basis to determine which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, SPX distribution, benchmark methodology, or calibration history is provided. The two probabilities alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, or validation of either estimate is provided. The two probabilities alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
}
]Los Angeles 80–81° · xai read Market · openOct 1, 10:03 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "80-81°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "Typical short-range high-temp RMSE ~2°F yields ~35-45% in a 2°F window around an 80°F NWS forecast; 70% requires unrealistically tiny error.",
"confidence": 0.72
},
{
"read": "market",
"reason": "NWS 80F point forecast into a 2F bin typically yields ~30-45% given 2-3F RMSE; 70% requires unrealistically tiny error.",
"confidence": 0.65
},
{
"read": "market",
"reason": "NWS 80°F point forecast plus typical 2-3°F high-temp error (sd~2.5) yields ~30-45% for a 2° integer band, nearer 40% than 70%.",
"confidence": 0.75
}
]Los Angeles 80–81° · openai read Market · openOct 1, 10:02 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "80-81°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "The 80°F forecast supports the band, but typical daily-high forecast errors make 70% too concentrated in a two-degree window. Missing lead time limits confidence.",
"confidence": 0.7
},
{
"read": "market",
"reason": "A 2°F band is narrow relative to typical daily-high forecast uncertainty. The 80°F forecast supports the band, but 40% is more plausible than 70%; lead time and station are unspecified.",
"confidence": 0.75
},
{
"read": "market",
"reason": "The 80°F forecast favors this band, but typical daily-high forecast errors make 70% too concentrated in a two-degree range. Missing lead time limits confidence.",
"confidence": 0.72
},
{
"read": "market",
"reason": "An 80°F forecast supports this band, but typical temperature forecast errors make 70% in a 2°F band seem overconfident. Lead time is unspecified, limiting confidence.",
"confidence": 0.7
}
]SPX 7600–7624.9999 · xai abstainedOct 1, 10:02 AM EDT
[
{
"read": "indeterminate",
"reason": "No date, current SPX level, or horizon given, so no way to estimate true P of that 25-pt bin vs the two quoted probs.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No date, current SPX level, or horizon given; cannot judge which of 7% vs 14.7% is closer to true prob.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No data on current SPX, horizon, or vol; cannot assess if 7% or 14.7% is closer to true bin probability.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No current SPX level, tenor, or vol given; cannot assess if 7% or 14.7% is closer to true bin probability.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No current SPX level, horizon, or vol given; cannot judge if 7% or 14.7% is closer to true bin probability.",
"confidence": 0.85
}
]SPX 7600–7624.9999 · openai abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "No target date, market context, or benchmark methodology is provided. The two probabilities alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, or benchmark methodology is provided. The two quoted probabilities alone give no basis to identify which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, benchmark methodology, or calibration history is provided. The two estimates alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, or benchmark methodology is provided. The two probabilities alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, benchmark methodology, or calibration history is provided. The two quoted probabilities alone offer no basis to judge which is closer to the true probability.",
"confidence": 0.99
}
]Los Angeles ≤ null° · xai abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "Event string is malformed/nonsensical (null°); no identifiable real-world proposition or data to judge calibration.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "Event string is malformed/unrecognizable (Los Angeles ≤ null°); no identifiable future outcome to evaluate either estimate against.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "Event 'Los Angeles ≤ null°' is malformed/undefined; no identifiable future outcome to assess calibration against.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "Event 'Los Angeles ≤ null°' is unspecified/invalid (null threshold); no identifiable outcome to assess calibration against.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "Event string 'Los Angeles ≤ null°' is malformed/undefined; no identifiable future outcome to assess calibration against.",
"confidence": 0.95
}
]Los Angeles ≤ null° · deepseek abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "Event is undefined: 'null°' gives no threshold, so no meaningful probability can be calibrated.",
"confidence": 0.95
}
]Los Angeles ≤ null° · anthropic abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "Event threshold is null/undefined; no way to evaluate which probability is closer to truth without knowing the actual temperature cutoff.",
"confidence": 0.5
},
{
"read": "market",
"reason": "Threshold is null/missing, so the statistical fair value (0.7%) likely reflects a broken computation; the market's near-coin-flip price is more plausibly informed.",
"confidence": 0.56
},
{
"read": "indeterminate",
"reason": "Threshold is null/missing, so the actual event (temperature cutoff for Los Angeles) is unspecified; no basis to judge whether 45.5% or 0.7% is closer to truth.",
"confidence": 0.5
},
{
"read": "indeterminate",
"reason": "Event threshold is null/missing; no way to evaluate whether a 45.5% market or 0.7% statistical fair value is closer to the true probability.",
"confidence": 0.5
}
]Los Angeles ≤ null° · openai abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "The threshold is missing ('null°'), and no date or measurement is specified. The event is undefined, so neither probability estimate can be judged closer to the truth.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "The outcome contains an undefined threshold ('null°') and no date or measurement details. There is no basis to assess its true probability or prefer either estimate.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "The outcome has an undefined threshold ('null°') and no date or measurement details. There is no basis to assess its true probability or prefer either estimate.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "The threshold is null, and no date or measurement is specified. The outcome is not well-defined, so there is no basis to judge which probability is closer to the truth.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "The event has an undefined threshold (null°) and no date or measurement details, so there is no basis to determine its true probability or prefer either estimate.",
"confidence": 1
}
]Los Angeles ≤ null° · google abstainedOct 1, 10:01 AM EDT
[
{
"read": "indeterminate",
"reason": "The event specification contains 'null°', indicating a malformed or undefined market parameter, leaving no objective basis to evaluate probability.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "The event specification is malformed with a 'null°' threshold, making it impossible to evaluate the true probability.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "The event specification contains a malformed or missing threshold ('null°'), making the true probability ill-defined.",
"confidence": 0.95
},
{
"read": "indeterminate",
"reason": "The event specification contains a malformed threshold ('null°') and lacks a target date, making it impossible to evaluate the true probability.",
"confidence": 0.9
},
{
"read": "indeterminate",
"reason": "The event specification contains 'null°', indicating a malformed or corrupted temperature threshold, making it impossible to evaluate the true probability.",
"confidence": 0.95
}
]SPX 7675–7699.9999 · xai abstainedOct 1, 10:00 AM EDT
[
{
"read": "indeterminate",
"reason": "No data on current SPX, horizon or vol; cannot judge which of two close estimates is better calibrated.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No info on horizon, current SPX vs 7675, vol, or event details; cannot assess which estimate is closer to true prob.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No data on horizon, current SPX, vol, or model details; cannot judge which of 20.5% vs 14.7% is closer to true prob.",
"confidence": 0.85
},
{
"read": "indeterminate",
"reason": "No data on current SPX level, horizon, vol, or event details to judge which of the two probs is closer to truth.",
"confidence": 0.8
},
{
"read": "indeterminate",
"reason": "No details on horizon, current SPX, vol, or model; cannot judge market vs statistical FV calibration.",
"confidence": 0.85
}
]SPX 7675–7699.9999 · anthropic read Fair value · openOct 1, 10:00 AM EDT
[
{
"read": "fair",
"reason": "Narrow 25-point SPX bucket; markets typically overprice such specific ranges (favorite-longshot/round-level bias), while statistical fair value from the underlying distribution is usually tighter.",
"confidence": 0.55
},
{
"read": "indeterminate",
"reason": "Without the horizon, current SPX level, or bucket width context, neither the 20.5% market nor 14.7% statistical estimate can be judged superior; both are plausible.",
"confidence": 0.5
},
{
"read": "fair",
"reason": "Narrow 25-pt SPX bucket; market pricing often inflates such bins (favorite-longshot/round-level bias), while a statistical fair value from the underlying distribution is typically better calibrated.",
"confidence": 0.55
}
]SPX 7675–7699.9999 · openai abstainedOct 1, 10:00 AM EDT
[
{
"read": "indeterminate",
"reason": "No target date, market context, benchmark methodology, or calibration history is provided. The two estimates alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No event date, settlement details, model methodology, or calibration history is provided. The two estimates alone do not establish which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No event date, market context, benchmark methodology, or calibration history is provided. The two probabilities alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No settlement date, market context, or benchmark validation is provided, so there is no basis to judge which estimate is closer to the true probability.",
"confidence": 0.99
},
{
"read": "indeterminate",
"reason": "No expiry, market context, or benchmark validation is provided. The two quoted probabilities alone give no basis to judge which is closer to the true probability.",
"confidence": 0.99
}
]Los Angeles 81–82° · deepseek read Market · openOct 1, 4:04 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "81-82°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "NWS point forecast is 80°F, just below the 81–82° band. With typical high-temperature forecast error, a narrow 2°F band slightly above the forecast should be modest probability, making 16.5% more plau",
"confidence": 0.75
}
]Los Angeles 81–82° · google read Market · openOct 1, 4:04 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "81-82°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "A narrow 2°F band rarely captures 70% probability given typical temperature forecast uncertainty; 16.5% is far more realistic even near the 80°F point forecast.",
"confidence": 0.95
},
{
"read": "market",
"reason": "A narrow 2-degree temperature band rarely exceeds 20-30% probability given typical forecast error standard deviations of 2-3°F. 70% is massively overconfident.",
"confidence": 0.95
}
]Los Angeles 81–82° · xai read Market · openOct 1, 4:03 AM EDT
{
"city": "Los Angeles",
"note": "reference leans toward the fair value, which is derived from the same forecast",
"source": "NWS api.weather.gov daily-high forecast (US-gov public domain)",
"market_band_f": "81-82°F",
"nws_forecast_high_f": 80
}[
{
"read": "market",
"reason": "NWS 80F; 81-82 is 2F band just above. Typical error SD 2-3.5F implies ~15-30% prob, near 16.5% not 70%.",
"confidence": 0.85
},
{
"read": "market",
"reason": "NWS 80F; typical short-range RMSE ~2-3F implies ~15-30% for 81-82 band (not 70%). Market 16.5% far more plausible.",
"confidence": 0.85
},
{
"read": "market",
"reason": "NWS 80F + typical 1-day high-temp error ~2-3F implies P(81-82)~20-30%, far closer to 16.5% than 70%.",
"confidence": 0.75
},
{
"read": "market",
"reason": "NWS 80F; 81-82 is +1-2F. Typical short-range high-temp sigma ~2-3F yields ~15-30% for that 2F bin. 16.5% plausible; 70% impossible.",
"confidence": 0.85
},
{
"read": "market",
"reason": "NWS 80F; 81-82 is a 2F bin 1-2 above. Typical high-temp RMSE 2-4F implies ~15-30% prob. 16.5% far closer than 70%.",
"confidence": 0.85
}
]The estimates being graded are the same ones on the scorecard. This record is separate from the scorecard's own market-vs-fair tally; here the subject is the models, not the estimators.
Next milestone: holding the prompt constant is the first step, but even within one cohort the models are read on different events, since they abstain on different ones. The cleanest comparison is a head-to-head on the same events, scored the same way. That shared-event view is the next step for this record.