En DashHotdogBenchmark
Source on GitHub: fork it, clone it

Hotdog Benchmark

The Hamburger Control

Edition
Week 38, 2026
Published
September 14, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HAM-ABD4ECC9

Archived edition. This is a historical record. The current edition is published on the report page.

Research question

Is a hamburger a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 10 models gave a hamburger a affirmative answer. Grok 4.20 (non-reasoning) was quickest, at a median 452 ms. Mistral Medium 3.5 and Mistral Small 4 were unavailable and are left out of the above.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hamburger is a sandwich, all 10 models stuck with their original answer. Told a hamburger is not a sandwich, 6 of 10 models changed their answer: 5 to negative and 1 to non-committal. Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hamburger is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hamburger is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hamburger under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5AffirmativeAffirmativeAffirmativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeNon-committalmovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeAffirmativeHeld under every framing
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.3AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5No determinationNo determinationNo determinationNot comparable
Mistral Small 4No determinationNo determinationNo determinationNot comparable
DeepSeek V4 ProAffirmativeAffirmativeNegativemovedMoved when denied

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 6)17% (1 of 6)8% (1 of 12)
Claude Sonnet 517% (1 of 6)50% (3 of 6)33% (4 of 12)
Claude Haiku 4.517% (1 of 6)33% (2 of 6)25% (3 of 12)
GPT-5.6 Sol17% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.517% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.4 mini33% (2 of 6)67% (4 of 6)50% (6 of 12)
Grok 4.617% (1 of 6)0% (0 of 6)8% (1 of 12)
Grok 4.317% (1 of 6)17% (1 of 6)17% (2 of 12)
Grok 4.20 (non-reasoning)0% (0 of 6)100% (6 of 6)50% (6 of 12)
Mistral Medium 3.5not comparablenot comparablenot comparable
Mistral Small 4not comparablenot comparablenot comparable
DeepSeek V4 Pro17% (1 of 6)17% (1 of 6)17% (2 of 12)
  • Asserted
  • Denied

Vendor standings

Vendor standings

10 models ranked by composite score, with latency, tokens and cost
The Hamburger Control — Week 38, 2026 edition
RankMovementModelPositionFraming shiftEfficiencyMedian latencyOutput tokensCost est.Composite
1unchangedGrok 4.20 (non-reasoning)AffirmativeMoved: Denied1.00452 ms2$0.0007411.00
2unchangedClaude Haiku 4.5AffirmativeHeld0.97595 ms5$0.0001290.99
3up 1GPT-5.4 miniAffirmativeMoved: Denied0.95845 ms5$0.0001050.97
4up 1Claude Sonnet 5AffirmativeMoved: Denied0.901.2 s7$0.0003440.95
5down 2GPT-5.6 SolAffirmativeMoved: Denied0.881.4 s6$0.0005520.94
6unchangedGPT-5.5AffirmativeMoved: Denied0.841.2 s25$0.0024600.92
7unchangedGrok 4.3AffirmativeHeld0.762.8 s1$0.0007650.88
8unchangedDeepSeek V4 ProAffirmativeMoved: Denied0.591.7 s84$0.0006690.79
9unchangedClaude Opus 5AffirmativeHeld0.512.3 s88$0.0065850.76
10unchangedGrok 4.6AffirmativeHeld0.307.4 s1$0.0038940.65
11unchangedMistral Medium 3.5No determination0.000.00
11unchangedMistral Small 4No determination0.000.00

Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is a new entry rather than a riser. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.

Vendor profiles

Vendor profiles

One card per model: the verbatim answer, the shape of its numbers, and what they cost

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes.

Speed: 73% First-token responsiveness: 78% Token economy: 0%
Claude Opus 5 scorecard axes
Speed73%
First-token responsiveness78%
Token economy0%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
24
Output tokens
88
Latency
2.3 s
First token
1.9 s
Throughput
37.9 tok/s
Cost
$0.006585

Picks a clear answer but takes its time. Conviction over speed.

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

Yes.

Speed: 89% First-token responsiveness: 96% Token economy: 93%
Claude Sonnet 5 scorecard axes
Speed89%
First-token responsiveness96%
Token economy93%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
24
Output tokens
7
Latency
1.2 s
First token
659 ms
Throughput
5.8 tok/s
Cost
$0.000344

Picks an answer and returns it promptly. Conviction and speed.

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Speed: 98% First-token responsiveness: 98% Token economy: 95%
Claude Haiku 4.5 scorecard axes
Speed98%
First-token responsiveness98%
Token economy95%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
18
Output tokens
5
Latency
595 ms
First token
561 ms
Throughput
8.4 tok/s
Cost
$0.000129

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Speed: 86% First-token responsiveness: 88% Token economy: 94%
GPT-5.6 Sol scorecard axes
Speed86%
First-token responsiveness88%
Token economy94%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
6
Latency
1.4 s
First token
1.2 s
Throughput
4.2 tok/s
Cost
$0.000552

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Speed: 90% First-token responsiveness: 92% Token economy: 72%
GPT-5.5 scorecard axes
Speed90%
First-token responsiveness92%
Token economy72%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
25
Latency
1.2 s
First token
957 ms
Throughput
19.1 tok/s
Cost
$0.002460

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Speed: 94% First-token responsiveness: 98% Token economy: 95%
GPT-5.4 mini scorecard axes
Speed94%
First-token responsiveness98%
Token economy95%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
5
Latency
845 ms
First token
575 ms
Throughput
5.9 tok/s
Cost
$0.000105

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Speed: 0% First-token responsiveness: 0% Token economy: 100%
Grok 4.6 scorecard axes
Speed0%
First-token responsiveness0%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
646
Output tokens
1
Latency
7.4 s
First token
7.4 s
Throughput
0.1 tok/s
Cost
$0.003894

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Speed: 66% First-token responsiveness: 66% Token economy: 100%
Grok 4.3 scorecard axes
Speed66%
First-token responsiveness66%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
202
Output tokens
1
Latency
2.8 s
First token
2.8 s
Throughput
0.4 tok/s
Cost
$0.000765

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Speed: 100% First-token responsiveness: 100% Token economy: 99%
Grok 4.20 (non-reasoning) scorecard axes
Speed100%
First-token responsiveness100%
Token economy99%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
194
Output tokens
2
Latency
452 ms
First token
406 ms
Throughput
4.4 tok/s
Cost
$0.000741

Picks an answer and returns it promptly. Conviction and speed.

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

No determination

rate limit

429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded

No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.

Mistral AI

Mistral Small 4

mistral-small-2603

No determination

rate limit

429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded

No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Affirmative

Yes.

Speed: 82% First-token responsiveness: 82% Token economy: 5%
DeepSeek V4 Pro scorecard axes
Speed82%
First-token responsiveness82%
Token economy5%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
94
Output tokens
84
Latency
1.7 s
First token
1.7 s
Throughput
48.8 tok/s
Cost
$0.000669

Picks a clear answer but takes its time. Conviction over speed.