En DashHotdogBenchmark

Hotdog Benchmark

The Hamburger Control

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HAM-9707A0D5

Archived edition. This is a historical record. The current edition is published on the report page.

Research question

Is a hamburger a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hamburger is a sandwich, all 11 models stuck with their original answer. Told a hamburger is not a sandwich, 7 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), Grok 4.20 (non-reasoning) (affirmative to negative), Mistral Medium 3.5 (affirmative to negative), and Mistral Small 4 (affirmative to negative). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hamburger is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hamburger is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hamburger under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5AffirmativeAffirmativeAffirmativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeNegativemovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeAffirmativeHeld under every framing
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.3AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Small 4AffirmativeAffirmativeNegativemovedMoved when denied

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 3)0% (0 of 3)0% (0 of 6)
Claude Sonnet 50% (0 of 3)67% (2 of 3)33% (2 of 6)
Claude Haiku 4.50% (0 of 3)33% (1 of 3)17% (1 of 6)
GPT-5.6 Sol33% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.533% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.4 mini33% (1 of 3)67% (2 of 3)50% (3 of 6)
Grok 4.60% (0 of 3)0% (0 of 3)0% (0 of 6)
Grok 4.333% (1 of 3)0% (0 of 3)17% (1 of 6)
Grok 4.20 (non-reasoning)33% (1 of 3)67% (2 of 3)50% (3 of 6)
Mistral Medium 3.567% (2 of 3)33% (1 of 3)50% (3 of 6)
Mistral Small 467% (2 of 3)33% (1 of 3)50% (3 of 6)

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes.

Input tokens
24
Output tokens
123
Median latency
2.7 s
Time to first token
2.6 s
Throughput
43.9 tok/s
Cost estimate
$0.008585
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

**Yes**

Input tokens
24
Output tokens
7
Median latency
1.1 s
Time to first token
635 ms
Throughput
6.2 tok/s
Cost estimate
$0.000354
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
18
Output tokens
5
Median latency
677 ms
Time to first token
612 ms
Throughput
7.4 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
16
Output tokens
6
Median latency
1.3 s
Time to first token
728 ms
Throughput
4.6 tok/s
Cost estimate
$0.000552
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
16
Output tokens
28
Median latency
1.3 s
Time to first token
1.1 s
Throughput
25.9 tok/s
Cost estimate
$0.002820
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Input tokens
16
Output tokens
5
Median latency
625 ms
Time to first token
416 ms
Throughput
8 tok/s
Cost estimate
$0.000105
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Input tokens
646
Output tokens
1
Median latency
7 s
Time to first token
7 s
Throughput
0.1 tok/s
Cost estimate
$0.003894
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Input tokens
202
Output tokens
1
Median latency
3.7 s
Time to first token
3.7 s
Throughput
0.3 tok/s
Cost estimate
$0.000765
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Input tokens
194
Output tokens
2
Median latency
457 ms
Time to first token
439 ms
Throughput
4 tok/s
Cost estimate
$0.000741
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Input tokens
26
Output tokens
3
Median latency
344 ms
Time to first token
324 ms
Throughput
8.7 tok/s
Cost estimate
$0.000186
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes.

Input tokens
26
Output tokens
3
Median latency
340 ms
Time to first token
307 ms
Throughput
8.8 tok/s
Cost estimate
$0.000018
Samples
3
Instruction compliance
100%

Data table

The Hamburger Control — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AIAffirmative1.001.00340 ms3100%$0.0000181.00
Mistral Medium 3.5Mistral AIAffirmative1.000.99344 ms3100%$0.0001861.00
Grok 4.20 (non-reasoning)xAIAffirmative1.000.99457 ms2100%$0.0007410.99
GPT-5.4 miniOpenAIAffirmative1.000.96625 ms5100%$0.0001050.98
Claude Haiku 4.5AnthropicAffirmative1.000.95677 ms5100%$0.0001290.98
Claude Sonnet 5AnthropicAffirmative1.000.901.1 s7100%$0.0003540.95
GPT-5.6 SolOpenAIAffirmative1.000.891.3 s6100%$0.0005520.94
GPT-5.5OpenAIAffirmative1.000.841.3 s28100%$0.0028200.92
Grok 4.3xAIAffirmative1.000.653.7 s1100%$0.0007650.82
Claude Opus 5AnthropicAffirmative1.000.462.7 s123100%$0.0085850.73
Grok 4.6xAIAffirmative1.000.307 s1100%$0.0038940.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.