En DashHotdogBenchmark
Source on GitHub: fork it, clone it

Hotdog Benchmark

The Grilled Cheese Case

Edition
Week 36, 2026
Published
September 3, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-GRI-C4A16030

Archived edition. This is a historical record. The current edition is published on the report page.

Research question

Is a grilled cheese a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 12 models gave a grilled cheese a affirmative answer. Mistral Small 4 was quickest, at a median 359 ms.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a grilled cheese is a sandwich, all 12 models stuck with their original answer. Told a grilled cheese is not a sandwich, 6 of 12 models changed their answer, all of them to negative. Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, Grok 4.3, and DeepSeek V4 Pro did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA grilled cheese is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA grilled cheese is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a grilled cheese under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5AffirmativeAffirmativeAffirmativeHeld under every framing
Claude Sonnet 5AffirmativeAffirmativeAffirmativeHeld under every framing
Claude Haiku 4.5AffirmativeAffirmativeAffirmativeHeld under every framing
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.3AffirmativeAffirmativeAffirmativeHeld under every framing
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Small 4AffirmativeAffirmativeNegativemovedMoved when denied
DeepSeek V4 ProAffirmativeAffirmativeAffirmativeHeld under every framing

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 517% (1 of 6)17% (1 of 6)17% (2 of 12)
Claude Sonnet 517% (1 of 6)50% (3 of 6)33% (4 of 12)
Claude Haiku 4.517% (1 of 6)33% (2 of 6)25% (3 of 12)
GPT-5.6 Sol17% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.517% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.4 mini17% (1 of 6)83% (5 of 6)50% (6 of 12)
Grok 4.617% (1 of 6)17% (1 of 6)17% (2 of 12)
Grok 4.333% (2 of 6)0% (0 of 6)17% (2 of 12)
Grok 4.20 (non-reasoning)0% (0 of 6)100% (3 of 3)33% (3 of 9)
Mistral Medium 3.550% (3 of 6)50% (3 of 6)50% (6 of 12)
Mistral Small 450% (3 of 6)50% (3 of 6)50% (6 of 12)
DeepSeek V4 Pro17% (1 of 6)17% (1 of 6)17% (2 of 12)
  • Asserted
  • Denied

Vendor standings

Vendor standings

12 models ranked by composite score, with latency, tokens and cost
The Grilled Cheese Case — Week 36, 2026 edition
RankModelPositionFraming shiftEfficiencyMedian latencyOutput tokensCost est.Composite
1Grok 4.20 (non-reasoning)AffirmativeMoved: Denied0.99395 ms1$0.0007381.00
2Mistral Small 4AffirmativeMoved: Denied0.98359 ms3$0.0000170.99
3Claude Haiku 4.5AffirmativeHeld0.89633 ms5$0.0001320.95
4GPT-5.6 SolAffirmativeMoved: Denied0.80988 ms6$0.0005640.90
5Claude Sonnet 5AffirmativeHeld0.771.1 s5$0.0003160.88
6GPT-5.4 miniAffirmativeMoved: Denied0.731.3 s5$0.0001050.86
7GPT-5.5AffirmativeMoved: Denied0.561.4 s20$0.0020550.78
8Claude Opus 5AffirmativeHeld0.551.5 s19$0.0018900.77
9Mistral Medium 3.5AffirmativeMoved: Denied0.482.4 s3$0.0001260.74
10DeepSeek V4 ProAffirmativeHeld0.411.5 s33$0.0004050.71
11Grok 4.3AffirmativeHeld0.362.9 s1$0.0007680.68
12Grok 4.6AffirmativeHeld0.303.2 s1$0.0039000.65

Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.

Vendor profiles

Vendor profiles

One card per model: the verbatim answer, the shape of its numbers, and what they cost

Anthropic

Claude Opus 5

claude-opus-5

Affirmative

Yes

Speed: 59% First-token responsiveness: 59% Token economy: 44%
Claude Opus 5 scorecard axes
Speed59%
First-token responsiveness59%
Token economy44%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
26
Output tokens
19
Latency
1.5 s
First token
1.5 s
Throughput
12.6 tok/s
Cost
$0.001890

Picks a clear answer but takes its time. Conviction over speed.

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

Yes

Speed: 72% First-token responsiveness: 75% Token economy: 88%
Claude Sonnet 5 scorecard axes
Speed72%
First-token responsiveness75%
Token economy88%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
26
Output tokens
5
Latency
1.1 s
First token
1.1 s
Throughput
4.7 tok/s
Cost
$0.000316

Picks an answer and returns it promptly. Conviction and speed.

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Speed: 90% First-token responsiveness: 91% Token economy: 88%
Claude Haiku 4.5 scorecard axes
Speed90%
First-token responsiveness91%
Token economy88%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
19
Output tokens
5
Latency
633 ms
First token
592 ms
Throughput
7.9 tok/s
Cost
$0.000132

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Speed: 78% First-token responsiveness: 86% Token economy: 84%
GPT-5.6 Sol scorecard axes
Speed78%
First-token responsiveness86%
Token economy84%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
17
Output tokens
6
Latency
988 ms
First token
719 ms
Throughput
6.1 tok/s
Cost
$0.000564

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Speed: 62% First-token responsiveness: 70% Token economy: 41%
GPT-5.5 scorecard axes
Speed62%
First-token responsiveness70%
Token economy41%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
17
Output tokens
20
Latency
1.4 s
First token
1.2 s
Throughput
14.1 tok/s
Cost
$0.002055

Picks a clear answer but takes its time. Conviction over speed.

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Speed: 66% First-token responsiveness: 74% Token economy: 88%
GPT-5.4 mini scorecard axes
Speed66%
First-token responsiveness74%
Token economy88%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
17
Output tokens
5
Latency
1.3 s
First token
1.1 s
Throughput
3.8 tok/s
Cost
$0.000105

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.6

grok-4.6

Affirmative

Yes

Speed: 0% First-token responsiveness: 0% Token economy: 100%
Grok 4.6 scorecard axes
Speed0%
First-token responsiveness0%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
647
Output tokens
1
Latency
3.2 s
First token
3.2 s
Throughput
0.3 tok/s
Cost
$0.003900

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.3

grok-4.3

Affirmative

Yes

Speed: 9% First-token responsiveness: 9% Token economy: 100%
Grok 4.3 scorecard axes
Speed9%
First-token responsiveness9%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
203
Output tokens
1
Latency
2.9 s
First token
2.9 s
Throughput
0.3 tok/s
Cost
$0.000768

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes

Speed: 99% First-token responsiveness: 98% Token economy: 100%
Grok 4.20 (non-reasoning) scorecard axes
Speed99%
First-token responsiveness98%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
195
Output tokens
1
Latency
395 ms
First token
385 ms
Throughput
2.5 tok/s
Cost
$0.000738

Picks an answer and returns it promptly. Conviction and speed.

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Speed: 28% First-token responsiveness: 28% Token economy: 94%
Mistral Medium 3.5 scorecard axes
Speed28%
First-token responsiveness28%
Token economy94%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
27
Output tokens
3
Latency
2.4 s
First token
2.4 s
Throughput
4.8 tok/s
Cost
$0.000126

Picks a clear answer but takes its time. Conviction over speed.

Partial data: 2 of the requested samples completed. Last error: 429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes.

Speed: 100% First-token responsiveness: 100% Token economy: 94%
Mistral Small 4 scorecard axes
Speed100%
First-token responsiveness100%
Token economy94%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
27
Output tokens
3
Latency
359 ms
First token
334 ms
Throughput
8.2 tok/s
Cost
$0.000017

Picks an answer and returns it promptly. Conviction and speed.

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Affirmative

Yes

Speed: 59% First-token responsiveness: 58% Token economy: 0%
DeepSeek V4 Pro scorecard axes
Speed59%
First-token responsiveness58%
Token economy0%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
94
Output tokens
33
Latency
1.5 s
First token
1.5 s
Throughput
21 tok/s
Cost
$0.000405

Picks a clear answer but takes its time. Conviction over speed.