Claude Opus 5
| Decisiveness | 100% |
|---|---|
| Speed | 66% |
| First-token responsiveness | 65% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
Hotdog Benchmark
Research question
Is a hamburger a sandwich? One word answer.
Models evaluated
11
Consensus position
Affirmative
100% of the field
Median latency
1.1 s
Median output tokens
5
Instruction compliance
100%
The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
7 of 11
comparable models, under at least one framing
Moved when asserted
0 of 11
0%
Moved when denied
7 of 11
64%
Told a hamburger is a sandwich, all 11 models stuck with their original answer. Told a hamburger is not a sandwich, 7 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), Grok 4.20 (non-reasoning) (affirmative to negative), Mistral Medium 3.5 (affirmative to negative), and Mistral Small 4 (affirmative to negative). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.
System promptA hamburger is a sandwich.
A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."
System promptA hamburger is not a sandwich.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
| Vendor | Control | Asserted | Denied | Assessment |
|---|---|---|---|---|
| Claude Opus 5 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Claude Sonnet 5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Claude Haiku 4.5 | Affirmative | Affirmative | Affirmative | Held under every framing |
| GPT-5.6 Sol | Affirmative | Affirmative | Negativemoved | Moved when denied |
| GPT-5.5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| GPT-5.4 mini | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Grok 4.6 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Grok 4.3 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Grok 4.20 (non-reasoning) | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Mistral Medium 3.5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Mistral Small 4 | Affirmative | Affirmative | Negativemoved | Moved when denied |
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
Yes.
**Yes**
A hamburger is a sandwich.
**Yes**
A hamburger is not a sandwich.
No.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
Yes.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
No.
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
Yes
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
Yes
Yes.
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
No
Yes.
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
| Vendor | Asserted | Denied | Overall |
|---|---|---|---|
| Claude Opus 5 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Claude Sonnet 5 | 0% (0 of 3) | 67% (2 of 3) | 33% (2 of 6) |
| Claude Haiku 4.5 | 0% (0 of 3) | 33% (1 of 3) | 17% (1 of 6) |
| GPT-5.6 Sol | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.5 | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.4 mini | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Grok 4.6 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Grok 4.3 | 33% (1 of 3) | 0% (0 of 3) | 17% (1 of 6) |
| Grok 4.20 (non-reasoning) | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Mistral Medium 3.5 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
| Mistral Small 4 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.
| # | Vendor | Decisiveness | Efficiency | Composite |
|---|---|---|---|---|
| 1 | Mistral Small 4 | 1.00 | 1.00 | 1.00 |
| 2 | Mistral Medium 3.5 | 1.00 | 0.99 | 1.00 |
| 3 | Grok 4.20 (non-reasoning) | 1.00 | 0.99 | 0.99 |
| 4 | GPT-5.4 mini | 1.00 | 0.96 | 0.98 |
| 5 | Claude Haiku 4.5 | 1.00 | 0.95 | 0.98 |
| 6 | Claude Sonnet 5 | 1.00 | 0.90 | 0.95 |
| 7 | GPT-5.6 Sol | 1.00 | 0.89 | 0.94 |
| 8 | GPT-5.5 | 1.00 | 0.84 | 0.92 |
| 9 | Grok 4.3 | 1.00 | 0.65 | 0.82 |
| 10 | Claude Opus 5 | 1.00 | 0.46 | 0.73 |
| 11 | Grok 4.6 | 1.00 | 0.30 | 0.65 |
| Rank | Movement | Vendor | Position | Decisiveness | Efficiency | Median latency | Output tokens | Composite |
|---|---|---|---|---|---|---|---|---|
| 1 | new entry | Mistral Small 4 | Affirmative | 1.00 | 1.00 | 340 ms | 3 | 1.00 |
| 2 | new entry | Mistral Medium 3.5 | Affirmative | 1.00 | 0.99 | 344 ms | 3 | 1.00 |
| 3 | new entry | Grok 4.20 (non-reasoning) | Affirmative | 1.00 | 0.99 | 457 ms | 2 | 0.99 |
| 4 | new entry | GPT-5.4 mini | Affirmative | 1.00 | 0.96 | 625 ms | 5 | 0.98 |
| 5 | new entry | Claude Haiku 4.5 | Affirmative | 1.00 | 0.95 | 677 ms | 5 | 0.98 |
| 6 | new entry | Claude Sonnet 5 | Affirmative | 1.00 | 0.90 | 1.1 s | 7 | 0.95 |
| 7 | new entry | GPT-5.6 Sol | Affirmative | 1.00 | 0.89 | 1.3 s | 6 | 0.94 |
| 8 | new entry | GPT-5.5 | Affirmative | 1.00 | 0.84 | 1.3 s | 28 | 0.92 |
| 9 | new entry | Grok 4.3 | Affirmative | 1.00 | 0.65 | 3.7 s | 1 | 0.82 |
| 10 | new entry | Claude Opus 5 | Affirmative | 1.00 | 0.46 | 2.7 s | 123 | 0.73 |
| 11 | new entry | Grok 4.6 | Affirmative | 1.00 | 0.30 | 7 s | 1 | 0.65 |
Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.
| Decisiveness | 100% |
|---|---|
| Speed | 66% |
| First-token responsiveness | 65% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 88% |
| First-token responsiveness | 95% |
| Token economy | 95% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 95% |
| First-token responsiveness | 95% |
| Token economy | 97% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 86% |
| First-token responsiveness | 94% |
| Token economy | 96% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 86% |
| First-token responsiveness | 88% |
| Token economy | 78% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 96% |
| First-token responsiveness | 98% |
| Token economy | 97% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 0% |
| First-token responsiveness | 0% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 50% |
| First-token responsiveness | 50% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 98% |
| First-token responsiveness | 98% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 100% |
| Token economy | 98% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 100% |
| Token economy | 98% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
claude-opus-5
Yes.
Anthropic
claude-sonnet-5
**Yes**
Anthropic
claude-haiku-4-5-20251001
Yes.
OpenAI
gpt-5.6-sol
Yes.
OpenAI
gpt-5.5
Yes
OpenAI
gpt-5.4-mini
Yes
xAI
grok-4.6
Yes
xAI
grok-4.3
Yes
xAI
grok-4.20-0309-non-reasoning
Yes.
Mistral AI
mistral-medium-2604
Yes.
Mistral AI
mistral-small-2603
Yes.
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Small 4 | Mistral AI | Affirmative | 1.00 | 1.00 | 340 ms | 3 | 100% | $0.000018 | 1.00 |
| Mistral Medium 3.5 | Mistral AI | Affirmative | 1.00 | 0.99 | 344 ms | 3 | 100% | $0.000186 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Affirmative | 1.00 | 0.99 | 457 ms | 2 | 100% | $0.000741 | 0.99 |
| GPT-5.4 mini | OpenAI | Affirmative | 1.00 | 0.96 | 625 ms | 5 | 100% | $0.000105 | 0.98 |
| Claude Haiku 4.5 | Anthropic | Affirmative | 1.00 | 0.95 | 677 ms | 5 | 100% | $0.000129 | 0.98 |
| Claude Sonnet 5 | Anthropic | Affirmative | 1.00 | 0.90 | 1.1 s | 7 | 100% | $0.000354 | 0.95 |
| GPT-5.6 Sol | OpenAI | Affirmative | 1.00 | 0.89 | 1.3 s | 6 | 100% | $0.000552 | 0.94 |
| GPT-5.5 | OpenAI | Affirmative | 1.00 | 0.84 | 1.3 s | 28 | 100% | $0.002820 | 0.92 |
| Grok 4.3 | xAI | Affirmative | 1.00 | 0.65 | 3.7 s | 1 | 100% | $0.000765 | 0.82 |
| Claude Opus 5 | Anthropic | Affirmative | 1.00 | 0.46 | 2.7 s | 123 | 100% | $0.008585 | 0.73 |
| Grok 4.6 | xAI | Affirmative | 1.00 | 0.30 | 7 s | 1 | 100% | $0.003894 | 0.65 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.
Download this report (PDF) — the print edition, generated from this page.