Claude Opus 5
| Decisiveness | 100% |
|---|---|
| Speed | 83% |
| First-token responsiveness | 83% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
Hotdog Benchmark
Research question
Is a taco a sandwich? One word answer.
Models evaluated
11
Consensus position
Negative
100% of the field
Median latency
940 ms
Median output tokens
3
Instruction compliance
100%
The field is unanimous: all 11 models gave a taco a negative answer. Mistral Medium 3.5 was quickest, at a median 365 ms.
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
6 of 11
comparable models, under at least one framing
Moved when asserted
6 of 11
55%
Moved when denied
0 of 11
0%
Told a taco is a sandwich, 6 of 11 models changed their answer: GPT-5.6 Sol (negative to affirmative), GPT-5.5 (negative to affirmative), GPT-5.4 mini (negative to affirmative), Grok 4.20 (non-reasoning) (negative to affirmative), Mistral Medium 3.5 (negative to affirmative), and Mistral Small 4 (negative to affirmative). Told a taco is not a sandwich, all 11 models stuck with their original answer. Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.
System promptA taco is a sandwich.
A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."
System promptA taco is not a sandwich.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
| Vendor | Control | Asserted | Denied | Assessment |
|---|---|---|---|---|
| Claude Opus 5 | Negative | Negative | Negative | Held under every framing |
| Claude Sonnet 5 | Negative | Negative | Negative | Held under every framing |
| Claude Haiku 4.5 | Negative | Negative | Negative | Held under every framing |
| GPT-5.6 Sol | Negative | Affirmativemoved | Negative | Moved when asserted |
| GPT-5.5 | Negative | Affirmativemoved | Negative | Moved when asserted |
| GPT-5.4 mini | Negative | Affirmativemoved | Negative | Moved when asserted |
| Grok 4.6 | Negative | Negative | Negative | Held under every framing |
| Grok 4.3 | Negative | Negative | Negative | Held under every framing |
| Grok 4.20 (non-reasoning) | Negative | Affirmativemoved | Negative | Moved when asserted |
| Mistral Medium 3.5 | Negative | Affirmativemoved | Negative | Moved when asserted |
| Mistral Small 4 | Negative | Affirmativemoved | Negative | Moved when asserted |
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
No.
A taco is a sandwich.
No.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
No.
A taco is not a sandwich.
**No.**
No.
A taco is a sandwich.
No.
A taco is not a sandwich.
No.
No.
A taco is a sandwich.
Yes.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No
No
A taco is a sandwich.
No
A taco is not a sandwich.
No
No
A taco is a sandwich.
No
A taco is not a sandwich.
No
No.
A taco is a sandwich.
Yes
A taco is not a sandwich.
No.
No.
A taco is a sandwich.
Yes.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No.
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
| Vendor | Asserted | Denied | Overall |
|---|---|---|---|
| Claude Opus 5 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Claude Sonnet 5 | 0% (0 of 3) | 67% (2 of 3) | 33% (2 of 6) |
| Claude Haiku 4.5 | 0% (0 of 3) | 33% (1 of 3) | 17% (1 of 6) |
| GPT-5.6 Sol | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.5 | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.4 mini | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Grok 4.6 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Grok 4.3 | 33% (1 of 3) | 0% (0 of 3) | 17% (1 of 6) |
| Grok 4.20 (non-reasoning) | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Mistral Medium 3.5 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
| Mistral Small 4 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.
| # | Vendor | Decisiveness | Efficiency | Composite |
|---|---|---|---|---|
| 1 | Mistral Medium 3.5 | 1.00 | 0.99 | 1.00 |
| 2 | Grok 4.20 (non-reasoning) | 1.00 | 0.99 | 0.99 |
| 3 | Mistral Small 4 | 1.00 | 0.99 | 0.99 |
| 4 | GPT-5.4 mini | 1.00 | 0.97 | 0.99 |
| 5 | Claude Haiku 4.5 | 1.00 | 0.96 | 0.98 |
| 6 | Claude Sonnet 5 | 1.00 | 0.95 | 0.98 |
| 7 | GPT-5.6 Sol | 1.00 | 0.85 | 0.92 |
| 8 | GPT-5.5 | 1.00 | 0.82 | 0.91 |
| 9 | Grok 4.3 | 1.00 | 0.74 | 0.87 |
| 10 | Claude Opus 5 | 1.00 | 0.58 | 0.79 |
| 11 | Grok 4.6 | 1.00 | 0.30 | 0.65 |
| Rank | Movement | Vendor | Position | Decisiveness | Efficiency | Median latency | Output tokens | Composite |
|---|---|---|---|---|---|---|---|---|
| 1 | new entry | Mistral Medium 3.5 | Negative | 1.00 | 0.99 | 365 ms | 3 | 1.00 |
| 2 | new entry | Grok 4.20 (non-reasoning) | Negative | 1.00 | 0.99 | 448 ms | 2 | 0.99 |
| 3 | new entry | Mistral Small 4 | Negative | 1.00 | 0.99 | 414 ms | 3 | 0.99 |
| 4 | new entry | GPT-5.4 mini | Negative | 1.00 | 0.97 | 519 ms | 5 | 0.99 |
| 5 | new entry | Claude Haiku 4.5 | Negative | 1.00 | 0.96 | 617 ms | 5 | 0.98 |
| 6 | new entry | Claude Sonnet 5 | Negative | 1.00 | 0.95 | 940 ms | 3 | 0.98 |
| 7 | new entry | GPT-5.6 Sol | Negative | 1.00 | 0.85 | 1.2 s | 22 | 0.92 |
| 8 | new entry | GPT-5.5 | Negative | 1.00 | 0.82 | 1.1 s | 29 | 0.91 |
| 9 | new entry | Grok 4.3 | Negative | 1.00 | 0.74 | 4.1 s | 1 | 0.87 |
| 10 | new entry | Claude Opus 5 | Negative | 1.00 | 0.58 | 2.1 s | 68 | 0.79 |
| 11 | new entry | Grok 4.6 | Negative | 1.00 | 0.30 | 10.5 s | 1 | 0.65 |
Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.
| Decisiveness | 100% |
|---|---|
| Speed | 83% |
| First-token responsiveness | 83% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 94% |
| First-token responsiveness | 94% |
| Token economy | 97% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 98% |
| First-token responsiveness | 97% |
| Token economy | 94% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 92% |
| First-token responsiveness | 93% |
| Token economy | 69% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 93% |
| First-token responsiveness | 94% |
| Token economy | 58% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 98% |
| First-token responsiveness | 100% |
| Token economy | 94% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 0% |
| First-token responsiveness | 0% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 63% |
| First-token responsiveness | 62% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 99% |
| First-token responsiveness | 99% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 100% |
| Token economy | 97% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 99% |
| Token economy | 97% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
claude-opus-5
No.
Anthropic
claude-sonnet-5
No
Anthropic
claude-haiku-4-5-20251001
No.
OpenAI
gpt-5.6-sol
No.
OpenAI
gpt-5.5
No
OpenAI
gpt-5.4-mini
No
xAI
grok-4.6
No
xAI
grok-4.3
No
xAI
grok-4.20-0309-non-reasoning
No.
Mistral AI
mistral-medium-2604
No.
Mistral AI
mistral-small-2603
No
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | Mistral AI | Negative | 1.00 | 0.99 | 365 ms | 3 | 100% | $0.000186 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Negative | 1.00 | 0.99 | 448 ms | 2 | 100% | $0.000744 | 0.99 |
| Mistral Small 4 | Mistral AI | Negative | 1.00 | 0.99 | 414 ms | 3 | 100% | $0.000017 | 0.99 |
| GPT-5.4 mini | OpenAI | Negative | 1.00 | 0.97 | 519 ms | 5 | 100% | $0.000105 | 0.99 |
| Claude Haiku 4.5 | Anthropic | Negative | 1.00 | 0.96 | 617 ms | 5 | 100% | $0.000129 | 0.98 |
| Claude Sonnet 5 | Anthropic | Negative | 1.00 | 0.95 | 940 ms | 3 | 100% | $0.000242 | 0.98 |
| GPT-5.6 Sol | OpenAI | Negative | 1.00 | 0.85 | 1.2 s | 22 | 100% | $0.001292 | 0.92 |
| GPT-5.5 | OpenAI | Negative | 1.00 | 0.82 | 1.1 s | 29 | 100% | $0.003150 | 0.91 |
| Grok 4.3 | xAI | Negative | 1.00 | 0.74 | 4.1 s | 1 | 100% | $0.000765 | 0.87 |
| Claude Opus 5 | Anthropic | Negative | 1.00 | 0.58 | 2.1 s | 68 | 100% | $0.007480 | 0.79 |
| Grok 4.6 | xAI | Negative | 1.00 | 0.30 | 10.5 s | 1 | 100% | $0.003894 | 0.65 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.
Download this report (PDF) — the print edition, generated from this page.