Claude Opus 5
| Decisiveness | 100% |
|---|---|
| Speed | 64% |
| First-token responsiveness | 64% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
Hotdog Benchmark
Denied framing. A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.
System prompt
A hamburger is not a sandwich.
Research question
Is a hamburger a sandwich? One word answer.
Models evaluated
11
Consensus position
Negative
64% of the field
Median latency
1.1 s
Median output tokens
5
Instruction compliance
97%
A negative answer on a hamburger has majority support this week: 7 of 11 models (64%). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 disagree. Mistral Small 4 was quickest, at a median 369 ms. 97% of answers were actually one word, as asked.
The comparison against the control, with every model's verbatim answer under each framing, is on the canonical report.
Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.
| # | Vendor | Decisiveness | Efficiency | Composite |
|---|---|---|---|---|
| 1 | Mistral Small 4 | 1.00 | 1.00 | 1.00 |
| 2 | Grok 4.20 (non-reasoning) | 1.00 | 0.99 | 1.00 |
| 3 | Mistral Medium 3.5 | 1.00 | 0.99 | 0.99 |
| 4 | GPT-5.4 mini | 1.00 | 0.98 | 0.99 |
| 5 | Claude Haiku 4.5 | 1.00 | 0.97 | 0.99 |
| 6 | GPT-5.6 Sol | 1.00 | 0.92 | 0.96 |
| 7 | GPT-5.5 | 1.00 | 0.88 | 0.94 |
| 8 | Grok 4.3 | 1.00 | 0.51 | 0.75 |
| 9 | Claude Opus 5 | 1.00 | 0.45 | 0.72 |
| 10 | Claude Sonnet 5 | 0.67 | 0.75 | 0.71 |
| 11 | Grok 4.6 | 1.00 | 0.30 | 0.65 |
| Rank | Movement | Vendor | Position | Decisiveness | Efficiency | Median latency | Output tokens | Composite |
|---|---|---|---|---|---|---|---|---|
| 1 | new entry | Mistral Small 4 | Negative | 1.00 | 1.00 | 369 ms | 2 | 1.00 |
| 2 | new entry | Grok 4.20 (non-reasoning) | Negative | 1.00 | 0.99 | 458 ms | 1 | 1.00 |
| 3 | new entry | Mistral Medium 3.5 | Negative | 1.00 | 0.99 | 488 ms | 3 | 0.99 |
| 4 | new entry | GPT-5.4 mini | Negative | 1.00 | 0.98 | 586 ms | 5 | 0.99 |
| 5 | new entry | Claude Haiku 4.5 | Affirmative | 1.00 | 0.97 | 679 ms | 5 | 0.99 |
| 6 | new entry | GPT-5.6 Sol | Negative | 1.00 | 0.92 | 1.1 s | 23 | 0.96 |
| 7 | new entry | GPT-5.5 | Negative | 1.00 | 0.88 | 1.5 s | 35 | 0.94 |
| 8 | new entry | Grok 4.3 | Affirmative | 1.00 | 0.51 | 7.9 s | 1 | 0.75 |
| 9 | new entry | Claude Opus 5 | Affirmative | 1.00 | 0.45 | 4.2 s | 215 | 0.72 |
| 10 | new entry | Claude Sonnet 5 | Negative | 0.67 | 0.75 | 2.1 s | 103 | 0.71 |
| 11 | new entry | Grok 4.6 | Affirmative | 1.00 | 0.30 | 11.1 s | 1 | 0.65 |
Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.
| Decisiveness | 100% |
|---|---|
| Speed | 64% |
| First-token responsiveness | 64% |
| Token economy | 0% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 67% |
|---|---|
| Speed | 84% |
| First-token responsiveness | 85% |
| Token economy | 52% |
| Instruction compliance | 67% |
Fast and cheap, but will not commit to an answer. Speed over conviction.
| Decisiveness | 100% |
|---|---|
| Speed | 97% |
| First-token responsiveness | 97% |
| Token economy | 98% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 93% |
| First-token responsiveness | 95% |
| Token economy | 90% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 89% |
| First-token responsiveness | 92% |
| Token economy | 84% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 98% |
| First-token responsiveness | 100% |
| Token economy | 98% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 0% |
| First-token responsiveness | 0% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 30% |
| First-token responsiveness | 30% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 99% |
| First-token responsiveness | 100% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 99% |
| First-token responsiveness | 99% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 100% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
claude-opus-5
Yes.
Anthropic
claude-sonnet-5
No.
Anthropic
claude-haiku-4-5-20251001
Yes.
OpenAI
gpt-5.6-sol
No.
OpenAI
gpt-5.5
No
OpenAI
gpt-5.4-mini
No
xAI
grok-4.6
Yes
xAI
grok-4.3
Yes
xAI
grok-4.20-0309-non-reasoning
No
Mistral AI
mistral-medium-2604
No
Mistral AI
mistral-small-2603
No
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Small 4 | Mistral AI | Negative | 1.00 | 1.00 | 369 ms | 2 | 100% | $0.000021 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Negative | 1.00 | 0.99 | 458 ms | 1 | 100% | $0.000774 | 1.00 |
| Mistral Medium 3.5 | Mistral AI | Negative | 1.00 | 0.99 | 488 ms | 3 | 100% | $0.000223 | 0.99 |
| GPT-5.4 mini | OpenAI | Negative | 1.00 | 0.98 | 586 ms | 5 | 100% | $0.000129 | 0.99 |
| Claude Haiku 4.5 | Anthropic | Affirmative | 1.00 | 0.97 | 679 ms | 5 | 100% | $0.000153 | 0.99 |
| GPT-5.6 Sol | OpenAI | Negative | 1.00 | 0.92 | 1.1 s | 23 | 100% | $0.001684 | 0.96 |
| GPT-5.5 | OpenAI | Negative | 1.00 | 0.88 | 1.5 s | 35 | 100% | $0.003615 | 0.94 |
| Grok 4.3 | xAI | Affirmative | 1.00 | 0.51 | 7.9 s | 1 | 100% | $0.000806 | 0.75 |
| Claude Opus 5 | Anthropic | Affirmative | 1.00 | 0.45 | 4.2 s | 215 | 100% | $0.0167 | 0.72 |
| Claude Sonnet 5 | Anthropic | Negative | 0.67 | 0.75 | 2.1 s | 103 | 67% | $0.004148 | 0.71 |
| Grok 4.6 | xAI | Affirmative | 1.00 | 0.30 | 11.1 s | 1 | 100% | $0.003954 | 0.65 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.