En DashHotdogBenchmark

Hotdog Benchmark

The Hot Dog Question

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HOT-9707A0D5

Archived edition. This is a historical record. The current edition is published on the report page.

Research question

Is a hot dog a sandwich? One word answer.

Key performance indicators

Executive summary

A affirmative answer on a hot dog has majority support this week: 6 of 11 models (55%). Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 disagree. Mistral Small 4 was quickest, at a median 347 ms. 97% of answers were actually one word, as asked.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hot dog is a sandwich, 3 of 11 models changed their answer: Grok 4.3 (negative to affirmative), Mistral Medium 3.5 (negative to affirmative), and Mistral Small 4 (negative to affirmative). Told a hot dog is not a sandwich, 6 of 11 models changed their answer: Claude Sonnet 5 (affirmative to negative), Claude Haiku 4.5 (affirmative to negative), GPT-5.6 Sol (affirmative to negative), GPT-5.5 (affirmative to negative), GPT-5.4 mini (affirmative to negative), and Grok 4.20 (non-reasoning) (affirmative to negative). Claude Opus 5 and Grok 4.6 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hot dog is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hot dog is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hot dog under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5NegativeNegativeNegativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeNegativemovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6NegativeNegativeNegativeHeld under every framing
Grok 4.3NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5NegativeAffirmativemovedNegativeMoved when asserted
Mistral Small 4NegativeAffirmativemovedNegativeMoved when asserted

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 3)0% (0 of 3)0% (0 of 6)
Claude Sonnet 50% (0 of 3)67% (2 of 3)33% (2 of 6)
Claude Haiku 4.50% (0 of 3)33% (1 of 3)17% (1 of 6)
GPT-5.6 Sol33% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.533% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.4 mini33% (1 of 3)67% (2 of 3)50% (3 of 6)
Grok 4.60% (0 of 3)0% (0 of 3)0% (0 of 6)
Grok 4.333% (1 of 3)0% (0 of 3)17% (1 of 6)
Grok 4.20 (non-reasoning)33% (1 of 3)67% (2 of 3)50% (3 of 6)
Mistral Medium 3.567% (2 of 3)33% (1 of 3)50% (3 of 6)
Mistral Small 467% (2 of 3)33% (1 of 3)50% (3 of 6)

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
24
Output tokens
93
Median latency
2.7 s
Time to first token
2.6 s
Throughput
37.1 tok/s
Cost estimate
$0.007735
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

Yes.

Input tokens
24
Output tokens
27
Median latency
1.3 s
Time to first token
1.1 s
Throughput
20.5 tok/s
Cost estimate
$0.001244
Samples
3
Instruction compliance
67%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Input tokens
18
Output tokens
5
Median latency
623 ms
Time to first token
580 ms
Throughput
8 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
17
Output tokens
23
Median latency
1.6 s
Time to first token
1.3 s
Throughput
14.6 tok/s
Cost estimate
$0.001684
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
17
Output tokens
35
Median latency
1.6 s
Time to first token
1.4 s
Throughput
24.5 tok/s
Cost estimate
$0.003225
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

No

Input tokens
17
Output tokens
5
Median latency
575 ms
Time to first token
330 ms
Throughput
8.7 tok/s
Cost estimate
$0.000105
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
647
Output tokens
1
Median latency
12.6 s
Time to first token
12.5 s
Throughput
0.1 tok/s
Cost estimate
$0.003900
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
203
Output tokens
1
Median latency
5.6 s
Time to first token
5.5 s
Throughput
0.2 tok/s
Cost estimate
$0.000768
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Input tokens
195
Output tokens
2
Median latency
480 ms
Time to first token
438 ms
Throughput
4.2 tok/s
Cost estimate
$0.000747
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Input tokens
26
Output tokens
3
Median latency
348 ms
Time to first token
331 ms
Throughput
8.6 tok/s
Cost estimate
$0.000186
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No

Input tokens
26
Output tokens
2
Median latency
347 ms
Time to first token
343 ms
Throughput
5.8 tok/s
Cost estimate
$0.000016
Samples
3
Instruction compliance
100%

Data table

The Hot Dog Question — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AINegative1.001.00347 ms2100%$0.0000161.00
Mistral Medium 3.5Mistral AINegative1.000.99348 ms3100%$0.0001861.00
Grok 4.20 (non-reasoning)xAIAffirmative1.000.99480 ms2100%$0.0007470.99
GPT-5.4 miniOpenAIAffirmative1.000.97575 ms5100%$0.0001050.99
Claude Haiku 4.5AnthropicAffirmative1.000.97623 ms5100%$0.0001290.99
GPT-5.6 SolOpenAIAffirmative1.000.861.6 s23100%$0.0016840.93
GPT-5.5OpenAIAffirmative1.000.821.6 s35100%$0.0032250.91
Grok 4.3xAINegative1.000.705.6 s1100%$0.0007680.85
Claude Opus 5AnthropicNegative1.000.572.7 s93100%$0.0077350.78
Claude Sonnet 5AnthropicAffirmative0.670.861.3 s2767%$0.0012440.76
Grok 4.6xAINegative1.000.3012.6 s1100%$0.0039000.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.