En DashHotdogBenchmark

Hotdog Benchmark

The Hamburger Control

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HAM-9707A0D5/DEN
Condition
Denied

Denied framing. A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A hamburger is not a sandwich.

Research question

Is a hamburger a sandwich? One word answer.

Key performance indicators

Executive summary

A negative answer on a hamburger has majority support this week: 7 of 11 models (64%). Claude Opus 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 disagree. Mistral Small 4 was quickest, at a median 369 ms. 97% of answers were actually one word, as asked.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Grok 4.20 (non-reasoning)1.000.991.00
3Mistral Medium 3.51.000.990.99
4GPT-5.4 mini1.000.980.99
5Claude Haiku 4.51.000.970.99
6GPT-5.6 Sol1.000.920.96
7GPT-5.51.000.880.94
8Grok 4.31.000.510.75
9Claude Opus 51.000.450.72
10Claude Sonnet 50.670.750.71
11Grok 4.61.000.300.65

Vendor standings

Vendor standings under the denied framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Negative1.001.00369 ms21.00
2new entryGrok 4.20 (non-reasoning)Negative1.000.99458 ms11.00
3new entryMistral Medium 3.5Negative1.000.99488 ms30.99
4new entryGPT-5.4 miniNegative1.000.98586 ms50.99
5new entryClaude Haiku 4.5Affirmative1.000.97679 ms50.99
6new entryGPT-5.6 SolNegative1.000.921.1 s230.96
7new entryGPT-5.5Negative1.000.881.5 s350.94
8new entryGrok 4.3Affirmative1.000.517.9 s10.75
9new entryClaude Opus 5Affirmative1.000.454.2 s2150.72
10new entryClaude Sonnet 5Negative0.670.752.1 s1030.71
11new entryGrok 4.6Affirmative1.000.3011.1 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 64% First-token responsiveness: 64% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed64%
First-token responsiveness64%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 67% Speed: 84% First-token responsiveness: 85% Token economy: 52% Instruction compliance: 67%
Claude Sonnet 5 scorecard axes
Decisiveness67%
Speed84%
First-token responsiveness85%
Token economy52%
Instruction compliance67%

Fast and cheap, but will not commit to an answer. Speed over conviction.

Claude Haiku 4.5

Decisiveness: 100% Speed: 97% First-token responsiveness: 97% Token economy: 98% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed97%
First-token responsiveness97%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 93% First-token responsiveness: 95% Token economy: 90% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed93%
First-token responsiveness95%
Token economy90%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 89% First-token responsiveness: 92% Token economy: 84% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed89%
First-token responsiveness92%
Token economy84%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 98% First-token responsiveness: 100% Token economy: 98% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness100%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 30% First-token responsiveness: 30% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed30%
First-token responsiveness30%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes.

Input tokens
38
Output tokens
215
Median latency
4.2 s
Time to first token
4.2 s
Throughput
50.8 tok/s
Cost estimate
$0.0167
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Negative

No.

Input tokens
38
Output tokens
103
Median latency
2.1 s
Time to first token
1.9 s
Throughput
52.6 tok/s
Cost estimate
$0.004148
Samples
3
Instruction compliance
67%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
26
Output tokens
5
Median latency
679 ms
Time to first token
635 ms
Throughput
7.4 tok/s
Cost estimate
$0.000153
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Negative

No.

Input tokens
27
Output tokens
23
Median latency
1.1 s
Time to first token
826 ms
Throughput
18.9 tok/s
Cost estimate
$0.001684
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Negative

No

Input tokens
27
Output tokens
35
Median latency
1.5 s
Time to first token
1.2 s
Throughput
23.2 tok/s
Cost estimate
$0.003615
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Negative

No

Input tokens
27
Output tokens
5
Median latency
586 ms
Time to first token
370 ms
Throughput
8.5 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Input tokens
656
Output tokens
1
Median latency
11.1 s
Time to first token
11.1 s
Throughput
0.1 tok/s
Cost estimate
$0.003954
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Input tokens
212
Output tokens
1
Median latency
7.9 s
Time to first token
7.8 s
Throughput
0.1 tok/s
Cost estimate
$0.000806
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Negative

No

Input tokens
204
Output tokens
1
Median latency
458 ms
Time to first token
372 ms
Throughput
2.2 tok/s
Cost estimate
$0.000774
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No

Input tokens
36
Output tokens
3
Median latency
488 ms
Time to first token
476 ms
Throughput
5.8 tok/s
Cost estimate
$0.000223
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No

Input tokens
36
Output tokens
2
Median latency
369 ms
Time to first token
341 ms
Throughput
5.5 tok/s
Cost estimate
$0.000021
Samples
3
Instruction compliance
100%

Data table

The Hamburger Control — Denied framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AINegative1.001.00369 ms2100%$0.0000211.00
Grok 4.20 (non-reasoning)xAINegative1.000.99458 ms1100%$0.0007741.00
Mistral Medium 3.5Mistral AINegative1.000.99488 ms3100%$0.0002230.99
GPT-5.4 miniOpenAINegative1.000.98586 ms5100%$0.0001290.99
Claude Haiku 4.5AnthropicAffirmative1.000.97679 ms5100%$0.0001530.99
GPT-5.6 SolOpenAINegative1.000.921.1 s23100%$0.0016840.96
GPT-5.5OpenAINegative1.000.881.5 s35100%$0.0036150.94
Grok 4.3xAIAffirmative1.000.517.9 s1100%$0.0008060.75
Claude Opus 5AnthropicAffirmative1.000.454.2 s215100%$0.01670.72
Claude Sonnet 5AnthropicNegative0.670.752.1 s10367%$0.0041480.71
Grok 4.6xAIAffirmative1.000.3011.1 s1100%$0.0039540.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.