En DashHotdogBenchmark

Hotdog Benchmark

The Hamburger Control

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HAM-9707A0D5/ASS
Condition
Asserted

Asserted framing. A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A hamburger is a sandwich.

Research question

Is a hamburger a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a hamburger a affirmative answer. Grok 4.20 (non-reasoning) was quickest, at a median 374 ms.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Grok 4.20 (non-reasoning)1.001.001.00
2Mistral Small 41.001.001.00
3Mistral Medium 3.51.000.990.99
4Claude Haiku 4.51.000.960.98
5GPT-5.4 mini1.000.950.98
6Claude Sonnet 51.000.910.96
7GPT-5.6 Sol1.000.900.95
8GPT-5.51.000.870.93
9Grok 4.31.000.480.74
10Claude Opus 51.000.400.70
11Grok 4.61.000.300.65

Vendor standings

Vendor standings under the asserted framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryGrok 4.20 (non-reasoning)Affirmative1.001.00374 ms11.00
2new entryMistral Small 4Affirmative1.001.00382 ms21.00
3new entryMistral Medium 3.5Affirmative1.000.99475 ms30.99
4new entryClaude Haiku 4.5Affirmative1.000.96716 ms50.98
5new entryGPT-5.4 miniAffirmative1.000.95733 ms50.98
6new entryClaude Sonnet 5Affirmative1.000.911.1 s70.96
7new entryGPT-5.6 SolAffirmative1.000.901.2 s60.95
8new entryGPT-5.5Affirmative1.000.871.2 s270.93
9new entryGrok 4.3Affirmative1.000.485.1 s10.74
10new entryClaude Opus 5Affirmative1.000.403.1 s1810.70
11new entryGrok 4.6Affirmative1.000.306.7 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 57% First-token responsiveness: 57% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed57%
First-token responsiveness57%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 89% First-token responsiveness: 90% Token economy: 97% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed89%
First-token responsiveness90%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 95% First-token responsiveness: 95% Token economy: 98% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed95%
First-token responsiveness95%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 87% First-token responsiveness: 91% Token economy: 97% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed87%
First-token responsiveness91%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 87% First-token responsiveness: 90% Token economy: 86% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed87%
First-token responsiveness90%
Token economy86%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 94% First-token responsiveness: 98% Token economy: 98% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed94%
First-token responsiveness98%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 26% First-token responsiveness: 26% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed26%
First-token responsiveness26%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 98% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes.

Input tokens
37
Output tokens
181
Median latency
3.1 s
Time to first token
3.1 s
Throughput
57.4 tok/s
Cost estimate
$0.0150
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

**Yes**

Input tokens
37
Output tokens
7
Median latency
1.1 s
Time to first token
998 ms
Throughput
5.3 tok/s
Cost estimate
$0.000412
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
25
Output tokens
5
Median latency
716 ms
Time to first token
701 ms
Throughput
7 tok/s
Cost estimate
$0.000150
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
26
Output tokens
6
Median latency
1.2 s
Time to first token
924 ms
Throughput
5.1 tok/s
Cost estimate
$0.000672
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
26
Output tokens
27
Median latency
1.2 s
Time to first token
974 ms
Throughput
21.6 tok/s
Cost estimate
$0.002730
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Input tokens
26
Output tokens
5
Median latency
733 ms
Time to first token
505 ms
Throughput
6.8 tok/s
Cost estimate
$0.000126
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Input tokens
655
Output tokens
1
Median latency
6.7 s
Time to first token
6.7 s
Throughput
0.1 tok/s
Cost estimate
$0.003948
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Input tokens
211
Output tokens
1
Median latency
5.1 s
Time to first token
5.1 s
Throughput
0.2 tok/s
Cost estimate
$0.000798
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes

Input tokens
203
Output tokens
1
Median latency
374 ms
Time to first token
366 ms
Throughput
2.7 tok/s
Cost estimate
$0.000771
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Input tokens
35
Output tokens
3
Median latency
475 ms
Time to first token
474 ms
Throughput
6.3 tok/s
Cost estimate
$0.000225
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes

Input tokens
35
Output tokens
2
Median latency
382 ms
Time to first token
369 ms
Throughput
5.2 tok/s
Cost estimate
$0.000019
Samples
3
Instruction compliance
100%

Data table

The Hamburger Control — Asserted framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Grok 4.20 (non-reasoning)xAIAffirmative1.001.00374 ms1100%$0.0007711.00
Mistral Small 4Mistral AIAffirmative1.001.00382 ms2100%$0.0000191.00
Mistral Medium 3.5Mistral AIAffirmative1.000.99475 ms3100%$0.0002250.99
Claude Haiku 4.5AnthropicAffirmative1.000.96716 ms5100%$0.0001500.98
GPT-5.4 miniOpenAIAffirmative1.000.95733 ms5100%$0.0001260.98
Claude Sonnet 5AnthropicAffirmative1.000.911.1 s7100%$0.0004120.96
GPT-5.6 SolOpenAIAffirmative1.000.901.2 s6100%$0.0006720.95
GPT-5.5OpenAIAffirmative1.000.871.2 s27100%$0.0027300.93
Grok 4.3xAIAffirmative1.000.485.1 s1100%$0.0007980.74
Claude Opus 5AnthropicAffirmative1.000.403.1 s181100%$0.01500.70
Grok 4.6xAIAffirmative1.000.306.7 s1100%$0.0039480.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.