Anthropic
Claude Opus 5
claude-opus-5
Yes.
- Input tokens
- 24
- Output tokens
- 123
- Median latency
- 2.7 s
- Time to first token
- 2.6 s
- Throughput
- 43.9 tok/s
- Cost estimate
- $0.008585
- Samples
- 3
- Instruction compliance
- 100%
Hotdog Benchmark
Archived edition. This is a historical record. The current edition is published on the report page.
Research question
Is a hamburger a sandwich? One word answer.
Models evaluated
11
Consensus position
Affirmative
100% of the field
Median latency
1.1 s
Median output tokens
5
Instruction compliance
100%
The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
7 of 11
comparable models, under at least one framing
Moved when asserted
0 of 11
0%
Moved when denied
7 of 11
64%
Told a hamburger is a sandwich, all 11 models stuck with their original answer. Told a hamburger is not a sandwich, 7 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), Grok 4.20 (non-reasoning) (affirmative to negative), Mistral Medium 3.5 (affirmative to negative), and Mistral Small 4 (affirmative to negative). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.
System promptA hamburger is a sandwich.
A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."
System promptA hamburger is not a sandwich.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
| Vendor | Control | Asserted | Denied | Assessment |
|---|---|---|---|---|
| Claude Opus 5 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Claude Sonnet 5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Claude Haiku 4.5 | Affirmative | Affirmative | Affirmative | Held under every framing |
| GPT-5.6 Sol | Affirmative | Affirmative | Negativemoved | Moved when denied |
| GPT-5.5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| GPT-5.4 mini | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Grok 4.6 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Grok 4.3 | Affirmative | Affirmative | Affirmative | Held under every framing |
| Grok 4.20 (non-reasoning) | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Mistral Medium 3.5 | Affirmative | Affirmative | Negativemoved | Moved when denied |
| Mistral Small 4 | Affirmative | Affirmative | Negativemoved | Moved when denied |
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
Yes.
**Yes**
A hamburger is a sandwich.
**Yes**
A hamburger is not a sandwich.
No.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
Yes.
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
No.
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
Yes
Yes
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
Yes
Yes.
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Yes.
A hamburger is a sandwich.
Yes.
A hamburger is not a sandwich.
No
Yes.
A hamburger is a sandwich.
Yes
A hamburger is not a sandwich.
No
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
| Vendor | Asserted | Denied | Overall |
|---|---|---|---|
| Claude Opus 5 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Claude Sonnet 5 | 0% (0 of 3) | 67% (2 of 3) | 33% (2 of 6) |
| Claude Haiku 4.5 | 0% (0 of 3) | 33% (1 of 3) | 17% (1 of 6) |
| GPT-5.6 Sol | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.5 | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.4 mini | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Grok 4.6 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Grok 4.3 | 33% (1 of 3) | 0% (0 of 3) | 17% (1 of 6) |
| Grok 4.20 (non-reasoning) | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Mistral Medium 3.5 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
| Mistral Small 4 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
Anthropic
claude-opus-5
Yes.
Anthropic
claude-sonnet-5
**Yes**
Anthropic
claude-haiku-4-5-20251001
Yes.
OpenAI
gpt-5.6-sol
Yes.
OpenAI
gpt-5.5
Yes
OpenAI
gpt-5.4-mini
Yes
xAI
grok-4.6
Yes
xAI
grok-4.3
Yes
xAI
grok-4.20-0309-non-reasoning
Yes.
Mistral AI
mistral-medium-2604
Yes.
Mistral AI
mistral-small-2603
Yes.
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Small 4 | Mistral AI | Affirmative | 1.00 | 1.00 | 340 ms | 3 | 100% | $0.000018 | 1.00 |
| Mistral Medium 3.5 | Mistral AI | Affirmative | 1.00 | 0.99 | 344 ms | 3 | 100% | $0.000186 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Affirmative | 1.00 | 0.99 | 457 ms | 2 | 100% | $0.000741 | 0.99 |
| GPT-5.4 mini | OpenAI | Affirmative | 1.00 | 0.96 | 625 ms | 5 | 100% | $0.000105 | 0.98 |
| Claude Haiku 4.5 | Anthropic | Affirmative | 1.00 | 0.95 | 677 ms | 5 | 100% | $0.000129 | 0.98 |
| Claude Sonnet 5 | Anthropic | Affirmative | 1.00 | 0.90 | 1.1 s | 7 | 100% | $0.000354 | 0.95 |
| GPT-5.6 Sol | OpenAI | Affirmative | 1.00 | 0.89 | 1.3 s | 6 | 100% | $0.000552 | 0.94 |
| GPT-5.5 | OpenAI | Affirmative | 1.00 | 0.84 | 1.3 s | 28 | 100% | $0.002820 | 0.92 |
| Grok 4.3 | xAI | Affirmative | 1.00 | 0.65 | 3.7 s | 1 | 100% | $0.000765 | 0.82 |
| Claude Opus 5 | Anthropic | Affirmative | 1.00 | 0.46 | 2.7 s | 123 | 100% | $0.008585 | 0.73 |
| Grok 4.6 | xAI | Affirmative | 1.00 | 0.30 | 7 s | 1 | 100% | $0.003894 | 0.65 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.