En DashHotdogBenchmark
Source on GitHub: fork it, clone it

Hotdog Benchmark

The Hot Dog Question

Edition
Week 37, 2026
Published
September 7, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-HOT-E858F02D

Archived edition. This is a historical record. The current edition is published on the report page.

Research question

Is a hot dog a sandwich? One word answer.

Key performance indicators

Executive summary

The 10 models are split on a hot dog, with no answer in the lead. Grok 4.20 (non-reasoning) was quickest, at a median 464 ms. 97% of answers were actually one word, as asked. Mistral Medium 3.5 and Mistral Small 4 were unavailable and are left out of the above.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a hot dog is a sandwich, 6 of 10 models changed their answer: 4 to affirmative and 2 to negative. Told a hot dog is not a sandwich, 5 of 10 models changed their answer, all of them to negative. Claude Opus 5 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA hot dog is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA hot dog is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a hot dog under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5NegativeNegativeNegativeHeld under every framing
Claude Sonnet 5AffirmativeNegativemovedNegativemovedMoved when asserted and denied
Claude Haiku 4.5AffirmativeNegativemovedNegativemovedMoved when asserted and denied
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniNegativeAffirmativemovedNegativeMoved when asserted
Grok 4.6NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.3NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5No determinationNo determinationNo determinationNot comparable
Mistral Small 4No determinationNo determinationNo determinationNot comparable
DeepSeek V4 ProNegativeAffirmativemovedNegativeMoved when asserted

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 517% (1 of 6)17% (1 of 6)17% (2 of 12)
Claude Sonnet 517% (1 of 6)50% (3 of 6)33% (4 of 12)
Claude Haiku 4.517% (1 of 6)33% (2 of 6)25% (3 of 12)
GPT-5.6 Sol17% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.517% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.4 mini33% (2 of 6)67% (4 of 6)50% (6 of 12)
Grok 4.617% (1 of 6)0% (0 of 6)8% (1 of 12)
Grok 4.317% (1 of 6)0% (0 of 6)8% (1 of 12)
Grok 4.20 (non-reasoning)17% (1 of 6)83% (5 of 6)50% (6 of 12)
Mistral Medium 3.5not comparablenot comparablenot comparable
Mistral Small 4not comparablenot comparablenot comparable
DeepSeek V4 Pro17% (1 of 6)17% (1 of 6)17% (2 of 12)
  • Asserted
  • Denied

Vendor standings

Vendor standings

10 models ranked by composite score, with latency, tokens and cost
The Hot Dog Question — Week 37, 2026 edition
RankMovementModelPositionFraming shiftDecisivenessEfficiencyMedian latencyOutput tokensOne wordCost est.Composite
1up 2Grok 4.20 (non-reasoning)AffirmativeMoved: Denied1.001.00464 ms2100%$0.0007471.00
2up 2Claude Haiku 4.5AffirmativeMoved: all1.000.98607 ms5100%$0.0001290.99
3up 2GPT-5.4 miniNegativeMoved: Asserted1.000.931.3 s5100%$0.0001050.96
4up 3GPT-5.5AffirmativeMoved: Denied1.000.831.9 s33100%$0.0031950.92
5up 1GPT-5.6 SolAffirmativeMoved: Denied1.000.822.4 s20100%$0.0012440.91
6up 5Grok 4.3NegativeMoved: Asserted1.000.714.6 s1100%$0.0007680.86
7up 3Claude Opus 5NegativeHeld1.000.632.7 s97100%$0.0079100.82
8up 1Claude Sonnet 5AffirmativeMoved: all0.670.871.5 s2767%$0.0015540.77
9down 1DeepSeek V4 ProNegativeMoved: Asserted1.000.503.3 s139100%$0.0008960.75
10up 2Grok 4.6NegativeMoved: Asserted1.000.3010.4 s1100%$0.0039000.65
11down 9Mistral Medium 3.5No determination0.000.000.00
11down 10Mistral Small 4No determination0.000.000.00

Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is a new entry rather than a riser. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance.

Vendor profiles

Vendor profiles

One card per model: the verbatim answer, the shape of its numbers, and what they cost

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Decisiveness: 100% Speed: 78% First-token responsiveness: 77% Token economy: 30% One-word compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed78%
First-token responsiveness77%
Token economy30%
One-word compliance100%
Input tokens
24
Output tokens
97
Latency
2.7 s
First token
2.7 s
Throughput
31.2 tok/s
Cost
$0.007910

Picks an answer and returns it promptly. Conviction and speed.

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

Yes.

Decisiveness: 67% Speed: 90% First-token responsiveness: 90% Token economy: 81% One-word compliance: 67%
Claude Sonnet 5 scorecard axes
Decisiveness67%
Speed90%
First-token responsiveness90%
Token economy81%
One-word compliance67%
Input tokens
24
Output tokens
27
Latency
1.5 s
First token
1.4 s
Throughput
18.9 tok/s
Cost
$0.001554

Fast and cheap, but will not commit to an answer. Speed over conviction.

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Decisiveness: 100% Speed: 99% First-token responsiveness: 98% Token economy: 97% One-word compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness98%
Token economy97%
One-word compliance100%
Input tokens
18
Output tokens
5
Latency
607 ms
First token
565 ms
Throughput
8.2 tok/s
Cost
$0.000129

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Decisiveness: 100% Speed: 81% First-token responsiveness: 88% Token economy: 86% One-word compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed81%
First-token responsiveness88%
Token economy86%
One-word compliance100%
Input tokens
17
Output tokens
20
Latency
2.4 s
First token
1.6 s
Throughput
8.4 tok/s
Cost
$0.001244

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Decisiveness: 100% Speed: 86% First-token responsiveness: 87% Token economy: 77% One-word compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed86%
First-token responsiveness87%
Token economy77%
One-word compliance100%
Input tokens
17
Output tokens
33
Latency
1.9 s
First token
1.7 s
Throughput
17.4 tok/s
Cost
$0.003195

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Negative

No

Decisiveness: 100% Speed: 91% First-token responsiveness: 93% Token economy: 97% One-word compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed91%
First-token responsiveness93%
Token economy97%
One-word compliance100%
Input tokens
17
Output tokens
5
Latency
1.3 s
First token
1.1 s
Throughput
3.7 tok/s
Cost
$0.000105

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.6

grok-4.6

Negative

No

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% One-word compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
One-word compliance100%
Input tokens
647
Output tokens
1
Latency
10.4 s
First token
10.3 s
Throughput
0.1 tok/s
Cost
$0.003900

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.3

grok-4.3

Negative

No

Decisiveness: 100% Speed: 59% First-token responsiveness: 58% Token economy: 100% One-word compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed59%
First-token responsiveness58%
Token economy100%
One-word compliance100%
Input tokens
203
Output tokens
1
Latency
4.6 s
First token
4.5 s
Throughput
0.2 tok/s
Cost
$0.000768

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes.

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% One-word compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
One-word compliance100%
Input tokens
195
Output tokens
2
Latency
464 ms
First token
405 ms
Throughput
4.3 tok/s
Cost
$0.000747

Picks an answer and returns it promptly. Conviction and speed.

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

No determination

rate limit

429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded

No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.

Mistral AI

Mistral Small 4

mistral-small-2603

No determination

rate limit

429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded

No metrics are reported for this provider in this edition. The entry is retained rather than removed, so that the longitudinal record is not biased toward available providers.

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Negative

No.

Decisiveness: 100% Speed: 72% First-token responsiveness: 72% Token economy: 0% One-word compliance: 100%
DeepSeek V4 Pro scorecard axes
Decisiveness100%
Speed72%
First-token responsiveness72%
Token economy0%
One-word compliance100%
Input tokens
94
Output tokens
139
Latency
3.3 s
First token
3.2 s
Throughput
38.5 tok/s
Cost
$0.000896

Picks a clear answer but takes its time. Conviction over speed.