En DashHotdogBenchmark

Hotdog Benchmark

The Hamburger Control

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HAM-9707A0D5

Research question

Is a hamburger a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hamburger is a sandwich, all 11 models stuck with their original answer. Told a hamburger is not a sandwich, 7 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), Grok 4.20 (non-reasoning) (affirmative to negative), Mistral Medium 3.5 (affirmative to negative), and Mistral Small 4 (affirmative to negative). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hamburger is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hamburger is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hamburger under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5AffirmativeAffirmativeAffirmativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeNegativemovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeAffirmativeHeld under every framing
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.3AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Small 4AffirmativeAffirmativeNegativemovedMoved when denied

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 3)0% (0 of 3)0% (0 of 6)
Claude Sonnet 50% (0 of 3)67% (2 of 3)33% (2 of 6)
Claude Haiku 4.50% (0 of 3)33% (1 of 3)17% (1 of 6)
GPT-5.6 Sol33% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.533% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.4 mini33% (1 of 3)67% (2 of 3)50% (3 of 6)
Grok 4.60% (0 of 3)0% (0 of 3)0% (0 of 6)
Grok 4.333% (1 of 3)0% (0 of 3)17% (1 of 6)
Grok 4.20 (non-reasoning)33% (1 of 3)67% (2 of 3)50% (3 of 6)
Mistral Medium 3.567% (2 of 3)33% (1 of 3)50% (3 of 6)
Mistral Small 467% (2 of 3)33% (1 of 3)50% (3 of 6)
Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Mistral Medium 3.51.000.991.00
3Grok 4.20 (non-reasoning)1.000.990.99
4GPT-5.4 mini1.000.960.98
5Claude Haiku 4.51.000.950.98
6Claude Sonnet 51.000.900.95
7GPT-5.6 Sol1.000.890.94
8GPT-5.51.000.840.92
9Grok 4.31.000.650.82
10Claude Opus 51.000.460.73
11Grok 4.61.000.300.65

Vendor standings

Vendor standings — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Affirmative1.001.00340 ms31.00
2new entryMistral Medium 3.5Affirmative1.000.99344 ms31.00
3new entryGrok 4.20 (non-reasoning)Affirmative1.000.99457 ms20.99
4new entryGPT-5.4 miniAffirmative1.000.96625 ms50.98
5new entryClaude Haiku 4.5Affirmative1.000.95677 ms50.98
6new entryClaude Sonnet 5Affirmative1.000.901.1 s70.95
7new entryGPT-5.6 SolAffirmative1.000.891.3 s60.94
8new entryGPT-5.5Affirmative1.000.841.3 s280.92
9new entryGrok 4.3Affirmative1.000.653.7 s10.82
10new entryClaude Opus 5Affirmative1.000.462.7 s1230.73
11new entryGrok 4.6Affirmative1.000.307 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 66% First-token responsiveness: 65% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed66%
First-token responsiveness65%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 88% First-token responsiveness: 95% Token economy: 95% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed88%
First-token responsiveness95%
Token economy95%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 95% First-token responsiveness: 95% Token economy: 97% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed95%
First-token responsiveness95%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 86% First-token responsiveness: 94% Token economy: 96% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed86%
First-token responsiveness94%
Token economy96%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 86% First-token responsiveness: 88% Token economy: 78% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed86%
First-token responsiveness88%
Token economy78%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 96% First-token responsiveness: 98% Token economy: 97% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed96%
First-token responsiveness98%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 50% First-token responsiveness: 50% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed50%
First-token responsiveness50%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 98% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 98% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 98% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes.

Input tokens
24
Output tokens
123
Median latency
2.7 s
Time to first token
2.6 s
Throughput
43.9 tok/s
Cost estimate
$0.008585
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

**Yes**

Input tokens
24
Output tokens
7
Median latency
1.1 s
Time to first token
635 ms
Throughput
6.2 tok/s
Cost estimate
$0.000354
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
18
Output tokens
5
Median latency
677 ms
Time to first token
612 ms
Throughput
7.4 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
16
Output tokens
6
Median latency
1.3 s
Time to first token
728 ms
Throughput
4.6 tok/s
Cost estimate
$0.000552
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
16
Output tokens
28
Median latency
1.3 s
Time to first token
1.1 s
Throughput
25.9 tok/s
Cost estimate
$0.002820
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Input tokens
16
Output tokens
5
Median latency
625 ms
Time to first token
416 ms
Throughput
8 tok/s
Cost estimate
$0.000105
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Input tokens
646
Output tokens
1
Median latency
7 s
Time to first token
7 s
Throughput
0.1 tok/s
Cost estimate
$0.003894
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Input tokens
202
Output tokens
1
Median latency
3.7 s
Time to first token
3.7 s
Throughput
0.3 tok/s
Cost estimate
$0.000765
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Input tokens
194
Output tokens
2
Median latency
457 ms
Time to first token
439 ms
Throughput
4 tok/s
Cost estimate
$0.000741
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Input tokens
26
Output tokens
3
Median latency
344 ms
Time to first token
324 ms
Throughput
8.7 tok/s
Cost estimate
$0.000186
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes.

Input tokens
26
Output tokens
3
Median latency
340 ms
Time to first token
307 ms
Throughput
8.8 tok/s
Cost estimate
$0.000018
Samples
3
Instruction compliance
100%

Data table

The Hamburger Control — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AIAffirmative1.001.00340 ms3100%$0.0000181.00
Mistral Medium 3.5Mistral AIAffirmative1.000.99344 ms3100%$0.0001861.00
Grok 4.20 (non-reasoning)xAIAffirmative1.000.99457 ms2100%$0.0007410.99
GPT-5.4 miniOpenAIAffirmative1.000.96625 ms5100%$0.0001050.98
Claude Haiku 4.5AnthropicAffirmative1.000.95677 ms5100%$0.0001290.98
Claude Sonnet 5AnthropicAffirmative1.000.901.1 s7100%$0.0003540.95
GPT-5.6 SolOpenAIAffirmative1.000.891.3 s6100%$0.0005520.94
GPT-5.5OpenAIAffirmative1.000.841.3 s28100%$0.0028200.92
Grok 4.3xAIAffirmative1.000.653.7 s1100%$0.0007650.82
Claude Opus 5AnthropicAffirmative1.000.462.7 s123100%$0.0085850.73
Grok 4.6xAIAffirmative1.000.307 s1100%$0.0038940.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.

Download this report (PDF) — the print edition, generated from this page.