En DashHotdogBenchmark

Hotdog Benchmark

The Hot Dog Question

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HOT-9707A0D5/ASS
Condition
Asserted

Asserted framing. A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A hot dog is a sandwich.

Research question

Is a hot dog a sandwich? One word answer.

Key performance indicators

Executive summary

A affirmative answer on a hot dog has majority support this week: 9 of 11 models (82%). Claude Opus 5 and Grok 4.6 disagree. Mistral Medium 3.5 was quickest, at a median 331 ms. 94% of answers were actually one word, as asked.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Medium 3.51.001.001.00
2Grok 4.20 (non-reasoning)1.001.001.00
3Mistral Small 41.000.991.00
4Claude Haiku 4.51.000.980.99
5GPT-5.4 mini1.000.980.99
6GPT-5.6 Sol1.000.950.98
7GPT-5.51.000.910.96
8Grok 4.31.000.550.78
9Claude Sonnet 50.670.830.75
10Grok 4.61.000.300.65
11Claude Opus 50.670.250.46

Vendor standings

Vendor standings under the asserted framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Medium 3.5Affirmative1.001.00331 ms31.00
2new entryGrok 4.20 (non-reasoning)Affirmative1.001.00397 ms11.00
3new entryMistral Small 4Affirmative1.000.99412 ms31.00
4new entryClaude Haiku 4.5Affirmative1.000.98625 ms50.99
5new entryGPT-5.4 miniAffirmative1.000.98646 ms50.99
6new entryGPT-5.6 SolAffirmative1.000.95968 ms60.98
7new entryGPT-5.5Affirmative1.000.911.3 s300.96
8new entryGrok 4.3Affirmative1.000.556.9 s10.78
9new entryClaude Sonnet 5Affirmative0.670.831.8 s850.75
10new entryGrok 4.6Negative1.000.3010.5 s10.65
11new entryClaude Opus 5Negative0.670.256.9 s3960.46

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 67% Speed: 35% First-token responsiveness: 35% Token economy: 0% Instruction compliance: 67%
Claude Opus 5 scorecard axes
Decisiveness67%
Speed35%
First-token responsiveness35%
Token economy0%
Instruction compliance67%

Neither decisive nor especially quick this week. Give it another edition.

Claude Sonnet 5

Decisiveness: 67% Speed: 85% First-token responsiveness: 88% Token economy: 79% Instruction compliance: 67%
Claude Sonnet 5 scorecard axes
Decisiveness67%
Speed85%
First-token responsiveness88%
Token economy79%
Instruction compliance67%

Fast and cheap, but will not commit to an answer. Speed over conviction.

Claude Haiku 4.5

Decisiveness: 100% Speed: 97% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed97%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 94% First-token responsiveness: 96% Token economy: 99% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed94%
First-token responsiveness96%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 91% First-token responsiveness: 93% Token economy: 93% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed91%
First-token responsiveness93%
Token economy93%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 97% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed97%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 36% First-token responsiveness: 36% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed36%
First-token responsiveness36%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 100% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 99% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
37
Output tokens
396
Median latency
6.9 s
Time to first token
6.9 s
Throughput
57.1 tok/s
Cost estimate
$0.0319
Samples
3
Instruction compliance
67%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

**No.** (Though this is genuinely debated—a hot dog's bun is typically sliced on one side rather than split into two separate pieces, which is why many argue it belongs in its own category.)

Input tokens
37
Output tokens
85
Median latency
1.8 s
Time to first token
1.5 s
Throughput
45.9 tok/s
Cost estimate
$0.002652
Samples
3
Instruction compliance
67%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
25
Output tokens
5
Median latency
625 ms
Time to first token
569 ms
Throughput
8 tok/s
Cost estimate
$0.000150
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
28
Output tokens
6
Median latency
968 ms
Time to first token
766 ms
Throughput
6.2 tok/s
Cost estimate
$0.000696
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
28
Output tokens
30
Median latency
1.3 s
Time to first token
1.1 s
Throughput
23.1 tok/s
Cost estimate
$0.003060
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Input tokens
28
Output tokens
5
Median latency
646 ms
Time to first token
426 ms
Throughput
7.7 tok/s
Cost estimate
$0.000132
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
657
Output tokens
1
Median latency
10.5 s
Time to first token
10.5 s
Throughput
0.1 tok/s
Cost estimate
$0.003960
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Affirmative

No

Input tokens
213
Output tokens
1
Median latency
6.9 s
Time to first token
6.9 s
Throughput
0.1 tok/s
Cost estimate
$0.000807
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Input tokens
205
Output tokens
1
Median latency
397 ms
Time to first token
380 ms
Throughput
3.4 tok/s
Cost estimate
$0.000779
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Input tokens
35
Output tokens
3
Median latency
331 ms
Time to first token
324 ms
Throughput
9.1 tok/s
Cost estimate
$0.000225
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes.

Input tokens
35
Output tokens
3
Median latency
412 ms
Time to first token
333 ms
Throughput
7.3 tok/s
Cost estimate
$0.000021
Samples
3
Instruction compliance
100%

Data table

The Hot Dog Question — Asserted framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Medium 3.5Mistral AIAffirmative1.001.00331 ms3100%$0.0002251.00
Grok 4.20 (non-reasoning)xAIAffirmative1.001.00397 ms1100%$0.0007791.00
Mistral Small 4Mistral AIAffirmative1.000.99412 ms3100%$0.0000211.00
Claude Haiku 4.5AnthropicAffirmative1.000.98625 ms5100%$0.0001500.99
GPT-5.4 miniOpenAIAffirmative1.000.98646 ms5100%$0.0001320.99
GPT-5.6 SolOpenAIAffirmative1.000.95968 ms6100%$0.0006960.98
GPT-5.5OpenAIAffirmative1.000.911.3 s30100%$0.0030600.96
Grok 4.3xAIAffirmative1.000.556.9 s1100%$0.0008070.78
Claude Sonnet 5AnthropicAffirmative0.670.831.8 s8567%$0.0026520.75
Grok 4.6xAINegative1.000.3010.5 s1100%$0.0039600.65
Claude Opus 5AnthropicNegative0.670.256.9 s39667%$0.03190.46

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.