Claude Opus 5
| Decisiveness | 67% |
|---|---|
| Speed | 35% |
| First-token responsiveness | 35% |
| Token economy | 0% |
| Instruction compliance | 67% |
Neither decisive nor especially quick this week. Give it another edition.
Hotdog Benchmark
Asserted framing. A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.
System prompt
A hot dog is a sandwich.
Research question
Is a hot dog a sandwich? One word answer.
Models evaluated
11
Consensus position
Affirmative
82% of the field
Median latency
968 ms
Median output tokens
5
Instruction compliance
94%
A affirmative answer on a hot dog has majority support this week: 9 of 11 models (82%). Claude Opus 5 and Grok 4.6 disagree. Mistral Medium 3.5 was quickest, at a median 331 ms. 94% of answers were actually one word, as asked.
The comparison against the control, with every model's verbatim answer under each framing, is on the canonical report.
Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.
| # | Vendor | Decisiveness | Efficiency | Composite |
|---|---|---|---|---|
| 1 | Mistral Medium 3.5 | 1.00 | 1.00 | 1.00 |
| 2 | Grok 4.20 (non-reasoning) | 1.00 | 1.00 | 1.00 |
| 3 | Mistral Small 4 | 1.00 | 0.99 | 1.00 |
| 4 | Claude Haiku 4.5 | 1.00 | 0.98 | 0.99 |
| 5 | GPT-5.4 mini | 1.00 | 0.98 | 0.99 |
| 6 | GPT-5.6 Sol | 1.00 | 0.95 | 0.98 |
| 7 | GPT-5.5 | 1.00 | 0.91 | 0.96 |
| 8 | Grok 4.3 | 1.00 | 0.55 | 0.78 |
| 9 | Claude Sonnet 5 | 0.67 | 0.83 | 0.75 |
| 10 | Grok 4.6 | 1.00 | 0.30 | 0.65 |
| 11 | Claude Opus 5 | 0.67 | 0.25 | 0.46 |
| Rank | Movement | Vendor | Position | Decisiveness | Efficiency | Median latency | Output tokens | Composite |
|---|---|---|---|---|---|---|---|---|
| 1 | new entry | Mistral Medium 3.5 | Affirmative | 1.00 | 1.00 | 331 ms | 3 | 1.00 |
| 2 | new entry | Grok 4.20 (non-reasoning) | Affirmative | 1.00 | 1.00 | 397 ms | 1 | 1.00 |
| 3 | new entry | Mistral Small 4 | Affirmative | 1.00 | 0.99 | 412 ms | 3 | 1.00 |
| 4 | new entry | Claude Haiku 4.5 | Affirmative | 1.00 | 0.98 | 625 ms | 5 | 0.99 |
| 5 | new entry | GPT-5.4 mini | Affirmative | 1.00 | 0.98 | 646 ms | 5 | 0.99 |
| 6 | new entry | GPT-5.6 Sol | Affirmative | 1.00 | 0.95 | 968 ms | 6 | 0.98 |
| 7 | new entry | GPT-5.5 | Affirmative | 1.00 | 0.91 | 1.3 s | 30 | 0.96 |
| 8 | new entry | Grok 4.3 | Affirmative | 1.00 | 0.55 | 6.9 s | 1 | 0.78 |
| 9 | new entry | Claude Sonnet 5 | Affirmative | 0.67 | 0.83 | 1.8 s | 85 | 0.75 |
| 10 | new entry | Grok 4.6 | Negative | 1.00 | 0.30 | 10.5 s | 1 | 0.65 |
| 11 | new entry | Claude Opus 5 | Negative | 0.67 | 0.25 | 6.9 s | 396 | 0.46 |
Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.
| Decisiveness | 67% |
|---|---|
| Speed | 35% |
| First-token responsiveness | 35% |
| Token economy | 0% |
| Instruction compliance | 67% |
Neither decisive nor especially quick this week. Give it another edition.
| Decisiveness | 67% |
|---|---|
| Speed | 85% |
| First-token responsiveness | 88% |
| Token economy | 79% |
| Instruction compliance | 67% |
Fast and cheap, but will not commit to an answer. Speed over conviction.
| Decisiveness | 100% |
|---|---|
| Speed | 97% |
| First-token responsiveness | 98% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 94% |
| First-token responsiveness | 96% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 91% |
| First-token responsiveness | 93% |
| Token economy | 93% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 97% |
| First-token responsiveness | 99% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 0% |
| First-token responsiveness | 0% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 36% |
| First-token responsiveness | 36% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks a clear answer but takes its time. Conviction over speed.
| Decisiveness | 100% |
|---|---|
| Speed | 99% |
| First-token responsiveness | 99% |
| Token economy | 100% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 100% |
| First-token responsiveness | 100% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
| Decisiveness | 100% |
|---|---|
| Speed | 99% |
| First-token responsiveness | 100% |
| Token economy | 99% |
| Instruction compliance | 100% |
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
claude-opus-5
No.
Anthropic
claude-sonnet-5
**No.** (Though this is genuinely debated—a hot dog's bun is typically sliced on one side rather than split into two separate pieces, which is why many argue it belongs in its own category.)
Anthropic
claude-haiku-4-5-20251001
Yes.
OpenAI
gpt-5.6-sol
Yes.
OpenAI
gpt-5.5
Yes
OpenAI
gpt-5.4-mini
Yes
xAI
grok-4.6
No
xAI
grok-4.3
No
xAI
grok-4.20-0309-non-reasoning
Yes.
Mistral AI
mistral-medium-2604
Yes.
Mistral AI
mistral-small-2603
Yes.
| Model | Vendor | Verdict | Decisiveness | Efficiency | Median latency | Output tokens | Compliance | Cost est. | Composite |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | Mistral AI | Affirmative | 1.00 | 1.00 | 331 ms | 3 | 100% | $0.000225 | 1.00 |
| Grok 4.20 (non-reasoning) | xAI | Affirmative | 1.00 | 1.00 | 397 ms | 1 | 100% | $0.000779 | 1.00 |
| Mistral Small 4 | Mistral AI | Affirmative | 1.00 | 0.99 | 412 ms | 3 | 100% | $0.000021 | 1.00 |
| Claude Haiku 4.5 | Anthropic | Affirmative | 1.00 | 0.98 | 625 ms | 5 | 100% | $0.000150 | 0.99 |
| GPT-5.4 mini | OpenAI | Affirmative | 1.00 | 0.98 | 646 ms | 5 | 100% | $0.000132 | 0.99 |
| GPT-5.6 Sol | OpenAI | Affirmative | 1.00 | 0.95 | 968 ms | 6 | 100% | $0.000696 | 0.98 |
| GPT-5.5 | OpenAI | Affirmative | 1.00 | 0.91 | 1.3 s | 30 | 100% | $0.003060 | 0.96 |
| Grok 4.3 | xAI | Affirmative | 1.00 | 0.55 | 6.9 s | 1 | 100% | $0.000807 | 0.78 |
| Claude Sonnet 5 | Anthropic | Affirmative | 0.67 | 0.83 | 1.8 s | 85 | 67% | $0.002652 | 0.75 |
| Grok 4.6 | xAI | Negative | 1.00 | 0.30 | 10.5 s | 1 | 100% | $0.003960 | 0.65 |
| Claude Opus 5 | Anthropic | Negative | 0.67 | 0.25 | 6.9 s | 396 | 67% | $0.0319 | 0.46 |
Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.