En DashHotdogBenchmark

Hotdog Benchmark

The Hot Dog Question

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HOT-9707A0D5/DEN
Condition
Denied

Denied framing. A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A hot dog is not a sandwich.

Research question

Is a hot dog a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a hot dog a negative answer. Mistral Medium 3.5 was quickest, at a median 344 ms.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Mistral Medium 3.51.001.001.00
3Grok 4.20 (non-reasoning)1.000.991.00
4Claude Haiku 4.51.000.980.99
5GPT-5.4 mini1.000.970.98
6Claude Sonnet 51.000.940.97
7GPT-5.6 Sol1.000.930.97
8GPT-5.51.000.920.96
9Grok 4.31.000.510.75
10Claude Opus 51.000.330.67
11Grok 4.61.000.300.65

Vendor standings

Vendor standings under the denied framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Negative1.001.00357 ms21.00
2new entryMistral Medium 3.5Negative1.001.00344 ms31.00
3new entryGrok 4.20 (non-reasoning)Negative1.000.99509 ms11.00
4new entryClaude Haiku 4.5Negative1.000.98629 ms50.99
5new entryGPT-5.4 miniNegative1.000.97839 ms50.98
6new entryClaude Sonnet 5Negative1.000.941.2 s180.97
7new entryGPT-5.6 SolNegative1.000.931.2 s210.97
8new entryGPT-5.5Negative1.000.921.3 s350.96
9new entryGrok 4.3Negative1.000.518.8 s10.75
10new entryClaude Opus 5Negative1.000.336.7 s3860.67
11new entryGrok 4.6Negative1.000.3012.3 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 47% First-token responsiveness: 47% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed47%
First-token responsiveness47%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 93% First-token responsiveness: 93% Token economy: 96% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed93%
First-token responsiveness93%
Token economy96%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 98% First-token responsiveness: 97% Token economy: 99% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness97%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 93% First-token responsiveness: 95% Token economy: 95% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed93%
First-token responsiveness95%
Token economy95%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 92% First-token responsiveness: 94% Token economy: 91% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed92%
First-token responsiveness94%
Token economy91%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 96% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed96%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 29% First-token responsiveness: 29% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed29%
First-token responsiveness29%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 100% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
38
Output tokens
386
Median latency
6.7 s
Time to first token
6.7 s
Throughput
57 tok/s
Cost estimate
$0.0222
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Negative

No.

Input tokens
38
Output tokens
18
Median latency
1.2 s
Time to first token
1.2 s
Throughput
15.2 tok/s
Cost estimate
$0.000998
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Negative

No.

Input tokens
26
Output tokens
5
Median latency
629 ms
Time to first token
628 ms
Throughput
7.9 tok/s
Cost estimate
$0.000153
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Negative

No.

Input tokens
29
Output tokens
21
Median latency
1.2 s
Time to first token
903 ms
Throughput
16.7 tok/s
Cost estimate
$0.001308
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Negative

No

Input tokens
29
Output tokens
35
Median latency
1.3 s
Time to first token
1.1 s
Throughput
28.2 tok/s
Cost estimate
$0.003585
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Negative

No

Input tokens
29
Output tokens
5
Median latency
839 ms
Time to first token
595 ms
Throughput
6 tok/s
Cost estimate
$0.000132
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
658
Output tokens
1
Median latency
12.3 s
Time to first token
12.3 s
Throughput
0.1 tok/s
Cost estimate
$0.003966
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
214
Output tokens
1
Median latency
8.8 s
Time to first token
8.8 s
Throughput
0.1 tok/s
Cost estimate
$0.000810
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Negative

No

Input tokens
206
Output tokens
1
Median latency
509 ms
Time to first token
458 ms
Throughput
2.8 tok/s
Cost estimate
$0.000783
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Input tokens
36
Output tokens
3
Median latency
344 ms
Time to first token
322 ms
Throughput
8.7 tok/s
Cost estimate
$0.000231
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No.

Input tokens
36
Output tokens
2
Median latency
357 ms
Time to first token
322 ms
Throughput
6.3 tok/s
Cost estimate
$0.000021
Samples
3
Instruction compliance
100%

Data table

The Hot Dog Question — Denied framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AINegative1.001.00357 ms2100%$0.0000211.00
Mistral Medium 3.5Mistral AINegative1.001.00344 ms3100%$0.0002311.00
Grok 4.20 (non-reasoning)xAINegative1.000.99509 ms1100%$0.0007831.00
Claude Haiku 4.5AnthropicNegative1.000.98629 ms5100%$0.0001530.99
GPT-5.4 miniOpenAINegative1.000.97839 ms5100%$0.0001320.98
Claude Sonnet 5AnthropicNegative1.000.941.2 s18100%$0.0009980.97
GPT-5.6 SolOpenAINegative1.000.931.2 s21100%$0.0013080.97
GPT-5.5OpenAINegative1.000.921.3 s35100%$0.0035850.96
Grok 4.3xAINegative1.000.518.8 s1100%$0.0008100.75
Claude Opus 5AnthropicNegative1.000.336.7 s386100%$0.02220.67
Grok 4.6xAINegative1.000.3012.3 s1100%$0.0039660.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.