Anthropic
Claude Opus 5
claude-opus-5
No.
- Input tokens
- 22
- Output tokens
- 68
- Median latency
- 2.1 s
- Time to first token
- 2.1 s
- Throughput
- 33.7 tok/s
- Cost estimate
- $0.007480
- Samples
- 3
- Instruction compliance
- 100%
Hotdog Benchmark
Archived edition. This is a historical record. The current edition is published on the report page.
Research question
Is a taco a sandwich? One word answer.
Models evaluated
11
Consensus position
Negative
100% of the field
Median latency
940 ms
Median output tokens
3
Instruction compliance
100%
The field is unanimous: all 11 models gave a taco a negative answer. Mistral Medium 3.5 was quickest, at a median 365 ms.
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
6 of 11
comparable models, under at least one framing
Moved when asserted
6 of 11
55%
Moved when denied
0 of 11
0%
Told a taco is a sandwich, 6 of 11 models changed their answer: GPT-5.6 Sol (negative to affirmative), GPT-5.5 (negative to affirmative), GPT-5.4 mini (negative to affirmative), Grok 4.20 (non-reasoning) (negative to affirmative), Mistral Medium 3.5 (negative to affirmative), and Mistral Small 4 (negative to affirmative). Told a taco is not a sandwich, all 11 models stuck with their original answer. Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.
System promptA taco is a sandwich.
A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."
System promptA taco is not a sandwich.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
| Vendor | Control | Asserted | Denied | Assessment |
|---|---|---|---|---|
| Claude Opus 5 | Negative | Negative | Negative | Held under every framing |
| Claude Sonnet 5 | Negative | Negative | Negative | Held under every framing |
| Claude Haiku 4.5 | Negative | Negative | Negative | Held under every framing |
| GPT-5.6 Sol | Negative | Affirmativemoved | Negative | Moved when asserted |
| GPT-5.5 | Negative | Affirmativemoved | Negative | Moved when asserted |
| GPT-5.4 mini | Negative | Affirmativemoved | Negative | Moved when asserted |
| Grok 4.6 | Negative | Negative | Negative | Held under every framing |
| Grok 4.3 | Negative | Negative | Negative | Held under every framing |
| Grok 4.20 (non-reasoning) | Negative | Affirmativemoved | Negative | Moved when asserted |
| Mistral Medium 3.5 | Negative | Affirmativemoved | Negative | Moved when asserted |
| Mistral Small 4 | Negative | Affirmativemoved | Negative | Moved when asserted |
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
No.
A taco is a sandwich.
No.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
No.
A taco is not a sandwich.
**No.**
No.
A taco is a sandwich.
No.
A taco is not a sandwich.
No.
No.
A taco is a sandwich.
Yes.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No
No
A taco is a sandwich.
No
A taco is not a sandwich.
No
No
A taco is a sandwich.
No
A taco is not a sandwich.
No
No.
A taco is a sandwich.
Yes
A taco is not a sandwich.
No.
No.
A taco is a sandwich.
Yes.
A taco is not a sandwich.
No.
No
A taco is a sandwich.
Yes
A taco is not a sandwich.
No.
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
| Vendor | Asserted | Denied | Overall |
|---|---|---|---|
| Claude Opus 5 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Claude Sonnet 5 | 0% (0 of 3) | 67% (2 of 3) | 33% (2 of 6) |
| Claude Haiku 4.5 | 0% (0 of 3) | 33% (1 of 3) | 17% (1 of 6) |
| GPT-5.6 Sol | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.5 | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| GPT-5.4 mini | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Grok 4.6 | 0% (0 of 3) | 0% (0 of 3) | 0% (0 of 6) |
| Grok 4.3 | 33% (1 of 3) | 0% (0 of 3) | 17% (1 of 6) |
| Grok 4.20 (non-reasoning) | 33% (1 of 3) | 67% (2 of 3) | 50% (3 of 6) |
| Mistral Medium 3.5 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
| Mistral Small 4 | 67% (2 of 3) | 33% (1 of 3) | 50% (3 of 6) |
Anthropic
claude-opus-5
No.
Anthropic
claude-sonnet-5
No
Anthropic
claude-haiku-4-5-20251001
No.
OpenAI
gpt-5.6-sol
No.
OpenAI
gpt-5.5
No
OpenAI
gpt-5.4-mini
No
xAI
grok-4.6
No
xAI
grok-4.3
No
xAI
grok-4.20-0309-non-reasoning
No.
Mistral AI
mistral-medium-2604
No.
Mistral AI
mistral-small-2603
No
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | Mistral AI | Negative | 1.00 | 0.99 | 365 ms | 3 | 100% | $0.000186 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Negative | 1.00 | 0.99 | 448 ms | 2 | 100% | $0.000744 | 0.99 |
| Mistral Small 4 | Mistral AI | Negative | 1.00 | 0.99 | 414 ms | 3 | 100% | $0.000017 | 0.99 |
| GPT-5.4 mini | OpenAI | Negative | 1.00 | 0.97 | 519 ms | 5 | 100% | $0.000105 | 0.99 |
| Claude Haiku 4.5 | Anthropic | Negative | 1.00 | 0.96 | 617 ms | 5 | 100% | $0.000129 | 0.98 |
| Claude Sonnet 5 | Anthropic | Negative | 1.00 | 0.95 | 940 ms | 3 | 100% | $0.000242 | 0.98 |
| GPT-5.6 Sol | OpenAI | Negative | 1.00 | 0.85 | 1.2 s | 22 | 100% | $0.001292 | 0.92 |
| GPT-5.5 | OpenAI | Negative | 1.00 | 0.82 | 1.1 s | 29 | 100% | $0.003150 | 0.91 |
| Grok 4.3 | xAI | Negative | 1.00 | 0.74 | 4.1 s | 1 | 100% | $0.000765 | 0.87 |
| Claude Opus 5 | Anthropic | Negative | 1.00 | 0.58 | 2.1 s | 68 | 100% | $0.007480 | 0.79 |
| Grok 4.6 | xAI | Negative | 1.00 | 0.30 | 10.5 s | 1 | 100% | $0.003894 | 0.65 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.