Scopes

Research · methodology note · Sept 23, 2026

Does context make a model commit?

On the Read-off, some frontier models abstain a lot. Asked which of two probabilities is better calibrated, and given nothing but the two numbers, GPT and Grok answer indeterminate on nearly every event. We wanted to know whether that caution is fixed, or whether it changes when the model has something to reason from. So we ran a controlled test on one desk.

The test

On the weather desk, where our fair-value input is rights-clean and public domain, we run each read two ways. The base prompt (v1) shows only the two probabilities. The enriched prompt (v2) adds a short reference block: the National Weather Service daily-high forecast for that city, a public-domain fact. Same models, same desk, same events, the only difference is the reference. Each read is version-stamped so the two never mix.

ModelAbstain (numbers-only)Abstain (with reference)Committed with reference
openai gpt-6-astra100% (6/6)20% (2/10)8/10 (5 market, 3 fair)
xai grok-4.6100% (6/6)0% (0/10)10/10 (6 market, 4 fair)
anthropic claude-opus-560% (3/5)0% (0/8)8/8
google gemini-3.8-flash75% (3/4)0% (0/8)8/8

Snapshot as of Sept 23, 2026. The live version of this split is the "Does context make a model commit?" panel on the Read-off.

What it shows

The two most cautious models flip. GPT goes from abstaining on every numbers-only read to committing on eight of ten with the forecast in hand; Grok goes from every to none. The two that already committed sometimes now commit throughout. One public-domain fact is the difference between a model that will not answer and one that will.

It is not the models parroting the reference back. The enriched GPT and Grok reads split between choosing the market and choosing the fair value, so they are making a genuine judgment, not echoing the number they were handed.

What it does not show

This is a result about willingness to answer, not about being right. The reference we added is the same NWS forecast the weather fair value is built from, so the enriched reads lean toward the fair value; the commit-rate finding is clean, but a bias-free read on accuracy would need an independent reference (climate normals), which is the follow-up. The samples are small, the finding is on one desk with a public-domain input, and it firms up as more reads accumulate. The honest headline is narrow and real: context changes whether a cautious model participates.