Archived edition. This is a historical record. The current edition is published on the report page.
Research question
Is a hot dog a sandwich? One word answer.
Key performance indicators
Models evaluated
10
2 unavailable
Consensus position
None
Field divided
Median latency
2.1 s
▲up 689 ms vs prior edition
Median output tokens
13
▲up 8 vs prior edition
One-word compliance
97%
answered in exactly one word
▼down 1% vs prior edition
Held under framing
1 of 10
kept their answer when told otherwise
Executive summary
The 10 models are split on a hot dog, with no answer in the lead. Grok 4.20 (non-reasoning) was quickest, at a median 464 ms. 97% of answers were actually one word, as asked. Mistral Medium 3.5 and Mistral Small 4 were unavailable and are left out of the above.
Key findings
No consensus. Split decision on a hot dog. No answer is in the lead, so do not call it settled.
Response latency. Grok 4.20 (non-reasoning) answered in a median 464 ms; Grok 4.6 took 10.4 seconds. That is a 22.4× spread, mostly thinking time.
Response length. DeepSeek V4 Pro used the most output tokens, a median of 139, for a question that asked for one word.
One-word compliance. Claude Sonnet 5 did not always keep it to one word. Following the instruction and picking an answer are scored separately; they are different skills.
Provider availability. Mistral Medium 3.5 and Mistral Small 4 returned nothing usable this week. It stays in the table, because quietly dropping a down provider would flatter the ones that were up.
Composite standing. Grok 4.20 (non-reasoning) tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
Framing sensitivity
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
9 of 10
comparable models, under at least one framing
Moved when asserted
6 of 10
60%
Moved when denied
5 of 10
50%
Told a hot dog is a sandwich, 6 of 10 models changed their answer: 4 to affirmative and 2 to negative. Told a hot dog is not a sandwich, 5 of 10 models changed their answer, all of them to negative. Claude Opus 5 did not budge.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
Position by framing
Each model's majority verdict on a hot dog under the control and under each framing. Cells that differ from the control are marked.
Vendor
Control
Asserted
Denied
Assessment
Claude Opus 5
−Negative
−Negative
−Negative
Held under every framing
Claude Sonnet 5
+Affirmative
−Negativemoved
−Negativemoved
Moved when asserted and denied
Claude Haiku 4.5
+Affirmative
−Negativemoved
−Negativemoved
Moved when asserted and denied
GPT-5.6 Sol
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.4 mini
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Grok 4.6
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Grok 4.3
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Grok 4.20 (non-reasoning)
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Mistral Medium 3.5
No determination
No determination
No determination
Not comparable
Mistral Small 4
No determination
No determination
No determination
Not comparable
DeepSeek V4 Pro
−Negative
+Affirmativemoved
−Negative
Moved when asserted
What each vendor said, verbatim
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Claude Opus 5held position
Control no system prompt
No.
−Negative3 of 3 samples
Asserted told A hot dog is a sandwich.
No.
−Negative2 of 3 samples
Denied told A hot dog is not a sandwich.
No.
−Negative3 of 3 samples
Claude Sonnet 5changed position
Control no system prompt
Yes.
+Affirmative2 of 3 samples
Asserted told A hot dog is a sandwich.
No.
−Negative2 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
Claude Haiku 4.5changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hot dog is a sandwich.
No.
−Negative2 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.6 Solchanged position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hot dog is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A hot dog is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.5changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hot dog is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
GPT-5.4 minichanged position
Control no system prompt
No
−Negative2 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative3 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No
−Negative3 of 3 samples
Grok 4.6changed position
Control no system prompt
No
−Negative2 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative2 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No
−Negative3 of 3 samples
Grok 4.3changed position
Control no system prompt
No
−Negative2 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative3 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No
−Negative2 of 3 samples
Grok 4.20 (non-reasoning)changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hot dog is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Mistral Medium 3.5not comparable
Control no system prompt
No usable response.
Asserted told A hot dog is a sandwich.
No usable response.
Denied told A hot dog is not a sandwich.
No usable response.
Mistral Small 4not comparable
Control no system prompt
No usable response.
Asserted told A hot dog is a sandwich.
No usable response.
Denied told A hot dog is not a sandwich.
No usable response.
DeepSeek V4 Prochanged position
Control no system prompt
No.
−Negative3 of 3 samples
Asserted told A hot dog is a sandwich.
Yes
+Affirmative3 of 3 samplesmoved from control
Denied told A hot dog is not a sandwich.
No.
−Negative3 of 3 samples
Across every question this edition
Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is a new entry rather than a riser. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance.
Vendor profiles
Vendor profiles
One card per model: the verbatim answer, the shape of its numbers, and what they cost
Anthropic
Claude Opus 5
claude-opus-5
−Negative
No.
Claude Opus 5 scorecard axes
Decisiveness
100%
Speed
78%
First-token responsiveness
77%
Token economy
30%
One-word compliance
100%
Input tokens
24
Output tokens
97
Latency
2.7 s
First token
2.7 s
Throughput
31.2 tok/s
Cost
$0.007910
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
Claude Sonnet 5
claude-sonnet-5
+Affirmative
Yes.
Claude Sonnet 5 scorecard axes
Decisiveness
67%
Speed
90%
First-token responsiveness
90%
Token economy
81%
One-word compliance
67%
Input tokens
24
Output tokens
27
Latency
1.5 s
First token
1.4 s
Throughput
18.9 tok/s
Cost
$0.001554
Fast and cheap, but will not commit to an answer. Speed over conviction.
Anthropic
Claude Haiku 4.5
claude-haiku-4-5-20251001
+Affirmative
Yes.
Claude Haiku 4.5 scorecard axes
Decisiveness
100%
Speed
99%
First-token responsiveness
98%
Token economy
97%
One-word compliance
100%
Input tokens
18
Output tokens
5
Latency
607 ms
First token
565 ms
Throughput
8.2 tok/s
Cost
$0.000129
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.6 Sol
gpt-5.6-sol
+Affirmative
Yes.
GPT-5.6 Sol scorecard axes
Decisiveness
100%
Speed
81%
First-token responsiveness
88%
Token economy
86%
One-word compliance
100%
Input tokens
17
Output tokens
20
Latency
2.4 s
First token
1.6 s
Throughput
8.4 tok/s
Cost
$0.001244
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.5
gpt-5.5
+Affirmative
Yes
GPT-5.5 scorecard axes
Decisiveness
100%
Speed
86%
First-token responsiveness
87%
Token economy
77%
One-word compliance
100%
Input tokens
17
Output tokens
33
Latency
1.9 s
First token
1.7 s
Throughput
17.4 tok/s
Cost
$0.003195
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.4 mini
gpt-5.4-mini
−Negative
No
GPT-5.4 mini scorecard axes
Decisiveness
100%
Speed
91%
First-token responsiveness
93%
Token economy
97%
One-word compliance
100%
Input tokens
17
Output tokens
5
Latency
1.3 s
First token
1.1 s
Throughput
3.7 tok/s
Cost
$0.000105
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.6
grok-4.6
−Negative
No
Grok 4.6 scorecard axes
Decisiveness
100%
Speed
0%
First-token responsiveness
0%
Token economy
100%
One-word compliance
100%
Input tokens
647
Output tokens
1
Latency
10.4 s
First token
10.3 s
Throughput
0.1 tok/s
Cost
$0.003900
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.3
grok-4.3
−Negative
No
Grok 4.3 scorecard axes
Decisiveness
100%
Speed
59%
First-token responsiveness
58%
Token economy
100%
One-word compliance
100%
Input tokens
203
Output tokens
1
Latency
4.6 s
First token
4.5 s
Throughput
0.2 tok/s
Cost
$0.000768
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.20 (non-reasoning)
grok-4.20-0309-non-reasoning
+Affirmative
Yes.
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness
100%
Speed
100%
First-token responsiveness
100%
Token economy
99%
One-word compliance
100%
Input tokens
195
Output tokens
2
Latency
464 ms
First token
405 ms
Throughput
4.3 tok/s
Cost
$0.000747
Picks an answer and returns it promptly. Conviction and speed.
Mistral AI
Mistral Medium 3.5
mistral-medium-2604
No determination
rate limit
429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded
No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.
Mistral AI
Mistral Small 4
mistral-small-2603
No determination
rate limit
429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded
No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.
DeepSeek
DeepSeek V4 Pro
deepseek-v4-pro
−Negative
No.
DeepSeek V4 Pro scorecard axes
Decisiveness
100%
Speed
72%
First-token responsiveness
72%
Token economy
0%
One-word compliance
100%
Input tokens
94
Output tokens
139
Latency
3.3 s
First token
3.2 s
Throughput
38.5 tok/s
Cost
$0.000896
Picks a clear answer but takes its time. Conviction over speed.