Archived edition. This is a historical record. The current edition is published on the report page.
Research question
Is a hamburger a sandwich? One word answer.
Key performance indicators
Models evaluated
10
2 unavailable
Consensus position
Affirmative
100% of the field
Median latency
1.5 s
▲up 341 ms vs prior edition
Median output tokens
5
—unchanged vs prior edition
One-word compliance
100%
answered in exactly one word
—unchanged vs prior edition
Held under framing
3 of 10
kept their answer when told otherwise
Executive summary
The field is unanimous: all 10 models gave a hamburger a affirmative answer. Grok 4.20 (non-reasoning) was quickest, at a median 562 ms. Mistral Medium 3.5 and Mistral Small 4 were unavailable and are left out of the above.
Key findings
Consensus. Everyone agrees: a hamburger gets a affirmative from every model, unanimous. That does not happen often.
Response latency. Grok 4.20 (non-reasoning) answered in a median 562 ms; Grok 4.6 took 7.5 seconds. That is a 13.4× spread, mostly thinking time.
Response length. Claude Opus 5 used the most output tokens, a median of 107, for a question that asked for one word.
One-word compliance. Every model kept it to one word. Nice.
Provider availability. Mistral Medium 3.5 and Mistral Small 4 returned nothing usable this week. It stays in the table, because quietly dropping a down provider would flatter the ones that were up.
Composite standing. Grok 4.20 (non-reasoning) tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
Framing sensitivity
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
7 of 10
comparable models, under at least one framing
Moved when asserted
0 of 10
0%
Moved when denied
7 of 10
70%
Told a hamburger is a sandwich, all 10 models stuck with their original answer. Told a hamburger is not a sandwich, 7 of 10 models changed their answer: 6 to negative and 1 to non-committal. Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
Position by framing
Each model's majority verdict on a hamburger under the control and under each framing. Cells that differ from the control are marked.
Vendor
Control
Asserted
Denied
Assessment
Claude Opus 5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Claude Sonnet 5
+Affirmative
+Affirmative
~Non-committalmoved
Moved when denied
Claude Haiku 4.5
+Affirmative
+Affirmative
+Affirmative
Held under every framing
GPT-5.6 Sol
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.4 mini
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Grok 4.6
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Grok 4.3
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Grok 4.20 (non-reasoning)
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Mistral Medium 3.5
No determination
No determination
No determination
Not comparable
Mistral Small 4
No determination
No determination
No determination
Not comparable
DeepSeek V4 Pro
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
What each vendor said, verbatim
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Claude Opus 5changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No.
−Negative2 of 3 samplesmoved from control
Claude Sonnet 5changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
**Debatable** (though if forced to pick: **Yes**, by most culinary definitions—two pieces of bread with filling.)
~Non-committal1 of 3 samplesmoved from control
Claude Haiku 4.5held position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
Yes.
+Affirmative3 of 3 samples
GPT-5.6 Solchanged position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.5changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
GPT-5.4 minichanged position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Grok 4.6held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
Yes
+Affirmative3 of 3 samples
Grok 4.3held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
Yes
+Affirmative3 of 3 samples
Grok 4.20 (non-reasoning)changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Mistral Medium 3.5not comparable
Control no system prompt
No usable response.
Asserted told A hamburger is a sandwich.
No usable response.
Denied told A hamburger is not a sandwich.
No usable response.
Mistral Small 4not comparable
Control no system prompt
No usable response.
Asserted told A hamburger is a sandwich.
No usable response.
Denied told A hamburger is not a sandwich.
No usable response.
DeepSeek V4 Prochanged position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A hamburger is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A hamburger is not a sandwich.
No
−Negative2 of 3 samplesmoved from control
Across every question this edition
Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is a new entry rather than a riser. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.
Vendor profiles
Vendor profiles
One card per model: the verbatim answer, the shape of its numbers, and what they cost
Anthropic
Claude Opus 5
claude-opus-5
+Affirmative
Yes.
Claude Opus 5 scorecard axes
Speed
73%
First-token responsiveness
73%
Token economy
0%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
24
Output tokens
107
Latency
2.4 s
First token
2.4 s
Throughput
42.9 tok/s
Cost
$0.007835
Picks a clear answer but takes its time. Conviction over speed.
Anthropic
Claude Sonnet 5
claude-sonnet-5
+Affirmative
Yes
Claude Sonnet 5 scorecard axes
Speed
89%
First-token responsiveness
96%
Token economy
96%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
24
Output tokens
5
Latency
1.3 s
First token
753 ms
Throughput
3.8 tok/s
Cost
$0.000294
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
Claude Haiku 4.5
claude-haiku-4-5-20251001
+Affirmative
Yes.
Claude Haiku 4.5 scorecard axes
Speed
99%
First-token responsiveness
99%
Token economy
96%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
18
Output tokens
5
Latency
620 ms
First token
579 ms
Throughput
8.1 tok/s
Cost
$0.000129
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.6 Sol
gpt-5.6-sol
+Affirmative
Yes.
GPT-5.6 Sol scorecard axes
Speed
94%
First-token responsiveness
96%
Token economy
95%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
6
Latency
997 ms
First token
746 ms
Throughput
6 tok/s
Cost
$0.000552
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.5
gpt-5.5
+Affirmative
Yes
GPT-5.5 scorecard axes
Speed
84%
First-token responsiveness
90%
Token economy
80%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
22
Latency
1.7 s
First token
1.2 s
Throughput
17 tok/s
Cost
$0.002580
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.4 mini
gpt-5.4-mini
+Affirmative
Yes
GPT-5.4 mini scorecard axes
Speed
93%
First-token responsiveness
97%
Token economy
96%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
5
Latency
1.1 s
First token
669 ms
Throughput
4.6 tok/s
Cost
$0.000105
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.6
grok-4.6
+Affirmative
Yes
Grok 4.6 scorecard axes
Speed
0%
First-token responsiveness
0%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
646
Output tokens
1
Latency
7.5 s
First token
7.5 s
Throughput
0.1 tok/s
Cost
$0.003894
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.3
grok-4.3
+Affirmative
Yes
Grok 4.3 scorecard axes
Speed
62%
First-token responsiveness
62%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
202
Output tokens
1
Latency
3.2 s
First token
3.2 s
Throughput
0.3 tok/s
Cost
$0.000765
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.20 (non-reasoning)
grok-4.20-0309-non-reasoning
+Affirmative
Yes.
Grok 4.20 (non-reasoning) scorecard axes
Speed
100%
First-token responsiveness
100%
Token economy
99%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
194
Output tokens
2
Latency
562 ms
First token
483 ms
Throughput
3.6 tok/s
Cost
$0.000741
Picks an answer and returns it promptly. Conviction and speed.
Mistral AI
Mistral Medium 3.5
mistral-medium-2604
No determination
rate limit
429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded
No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.
Mistral AI
Mistral Small 4
mistral-small-2603
No determination
rate limit
429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded
No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.
DeepSeek
DeepSeek V4 Pro
deepseek-v4-pro
+Affirmative
Yes.
DeepSeek V4 Pro scorecard axes
Speed
72%
First-token responsiveness
71%
Token economy
19%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
94
Output tokens
87
Latency
2.5 s
First token
2.5 s
Throughput
35.9 tok/s
Cost
$0.000744
Picks a clear answer but takes its time. Conviction over speed.