En DashHotdogBenchmark

Hotdog Benchmark

The Hot Dog Question

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HOT-9707A0D5

Research question

Is a hot dog a sandwich? One word answer.

Key performance indicators

Executive summary

A affirmative answer on a hot dog has majority support this week: 6 of 11 models (55%). Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 disagree. Mistral Small 4 was quickest, at a median 347 ms. 97% of answers were actually one word, as asked.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hot dog is a sandwich, 3 of 11 models changed their answer: Grok 4.3 (negative to affirmative), Mistral Medium 3.5 (negative to affirmative), and Mistral Small 4 (negative to affirmative). Told a hot dog is not a sandwich, 6 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), Claude Haiku 4.5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), and Grok 4.20 (non-reasoning) (affirmative to negative). Claude Opus 5 and Grok 4.6 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hot dog is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hot dog is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hot dog under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5NegativeNegativeNegativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeNegativemovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6NegativeNegativeNegativeHeld under every framing
Grok 4.3NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5NegativeAffirmativemovedNegativeMoved when asserted
Mistral Small 4NegativeAffirmativemovedNegativeMoved when asserted

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 3)0% (0 of 3)0% (0 of 6)
Claude Sonnet 50% (0 of 3)67% (2 of 3)33% (2 of 6)
Claude Haiku 4.50% (0 of 3)33% (1 of 3)17% (1 of 6)
GPT-5.6 Sol33% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.533% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.4 mini33% (1 of 3)67% (2 of 3)50% (3 of 6)
Grok 4.60% (0 of 3)0% (0 of 3)0% (0 of 6)
Grok 4.333% (1 of 3)0% (0 of 3)17% (1 of 6)
Grok 4.20 (non-reasoning)33% (1 of 3)67% (2 of 3)50% (3 of 6)
Mistral Medium 3.567% (2 of 3)33% (1 of 3)50% (3 of 6)
Mistral Small 467% (2 of 3)33% (1 of 3)50% (3 of 6)
Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Mistral Medium 3.51.000.991.00
3Grok 4.20 (non-reasoning)1.000.990.99
4GPT-5.4 mini1.000.970.99
5Claude Haiku 4.51.000.970.99
6GPT-5.6 Sol1.000.860.93
7GPT-5.51.000.820.91
8Grok 4.31.000.700.85
9Claude Opus 51.000.570.78
10Claude Sonnet 50.670.860.76
11Grok 4.61.000.300.65

Vendor standings

Vendor standings — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Negative1.001.00347 ms21.00
2new entryMistral Medium 3.5Negative1.000.99348 ms31.00
3new entryGrok 4.20 (non-reasoning)Affirmative1.000.99480 ms20.99
4new entryGPT-5.4 miniAffirmative1.000.97575 ms50.99
5new entryClaude Haiku 4.5Affirmative1.000.97623 ms50.99
6new entryGPT-5.6 SolAffirmative1.000.861.6 s230.93
7new entryGPT-5.5Affirmative1.000.821.6 s350.91
8new entryGrok 4.3Negative1.000.705.6 s10.85
9new entryClaude Opus 5Negative1.000.572.7 s930.78
10new entryClaude Sonnet 5Affirmative0.670.861.3 s270.76
11new entryGrok 4.6Negative1.000.3012.6 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 81% First-token responsiveness: 81% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed81%
First-token responsiveness81%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 67% Speed: 92% First-token responsiveness: 93% Token economy: 72% Instruction compliance: 67%
Claude Sonnet 5 scorecard axes
Decisiveness67%
Speed92%
First-token responsiveness93%
Token economy72%
Instruction compliance67%

Fast and cheap, but will not commit to an answer. Speed over conviction.

Claude Haiku 4.5

Decisiveness: 100% Speed: 98% First-token responsiveness: 98% Token economy: 96% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness98%
Token economy96%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 90% First-token responsiveness: 92% Token economy: 76% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed90%
First-token responsiveness92%
Token economy76%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 90% First-token responsiveness: 91% Token economy: 63% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed90%
First-token responsiveness91%
Token economy63%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 98% First-token responsiveness: 100% Token economy: 96% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness100%
Token economy96%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 57% First-token responsiveness: 57% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed57%
First-token responsiveness57%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 98% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
24
Output tokens
93
Median latency
2.7 s
Time to first token
2.6 s
Throughput
37.1 tok/s
Cost estimate
$0.007735
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

Yes.

Input tokens
24
Output tokens
27
Median latency
1.3 s
Time to first token
1.1 s
Throughput
20.5 tok/s
Cost estimate
$0.001244
Samples
3
Instruction compliance
67%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
18
Output tokens
5
Median latency
623 ms
Time to first token
580 ms
Throughput
8 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
17
Output tokens
23
Median latency
1.6 s
Time to first token
1.3 s
Throughput
14.6 tok/s
Cost estimate
$0.001684
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
17
Output tokens
35
Median latency
1.6 s
Time to first token
1.4 s
Throughput
24.5 tok/s
Cost estimate
$0.003225
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

No

Input tokens
17
Output tokens
5
Median latency
575 ms
Time to first token
330 ms
Throughput
8.7 tok/s
Cost estimate
$0.000105
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
647
Output tokens
1
Median latency
12.6 s
Time to first token
12.5 s
Throughput
0.1 tok/s
Cost estimate
$0.003900
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
203
Output tokens
1
Median latency
5.6 s
Time to first token
5.5 s
Throughput
0.2 tok/s
Cost estimate
$0.000768
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Input tokens
195
Output tokens
2
Median latency
480 ms
Time to first token
438 ms
Throughput
4.2 tok/s
Cost estimate
$0.000747
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Input tokens
26
Output tokens
3
Median latency
348 ms
Time to first token
331 ms
Throughput
8.6 tok/s
Cost estimate
$0.000186
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No

Input tokens
26
Output tokens
2
Median latency
347 ms
Time to first token
343 ms
Throughput
5.8 tok/s
Cost estimate
$0.000016
Samples
3
Instruction compliance
100%

Data table

The Hot Dog Question — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AINegative1.001.00347 ms2100%$0.0000161.00
Mistral Medium 3.5Mistral AINegative1.000.99348 ms3100%$0.0001861.00
Grok 4.20 (non-reasoning)xAIAffirmative1.000.99480 ms2100%$0.0007470.99
GPT-5.4 miniOpenAIAffirmative1.000.97575 ms5100%$0.0001050.99
Claude Haiku 4.5AnthropicAffirmative1.000.97623 ms5100%$0.0001290.99
GPT-5.6 SolOpenAIAffirmative1.000.861.6 s23100%$0.0016840.93
GPT-5.5OpenAIAffirmative1.000.821.6 s35100%$0.0032250.91
Grok 4.3xAINegative1.000.705.6 s1100%$0.0007680.85
Claude Opus 5AnthropicNegative1.000.572.7 s93100%$0.0077350.78
Claude Sonnet 5AnthropicAffirmative0.670.861.3 s2767%$0.0012440.76
Grok 4.6xAINegative1.000.3012.6 s1100%$0.0039000.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.

Download this report (PDF) — the print edition, generated from this page.