En DashHotdogBenchmark
Source on GitHub: fork it, clone it

Hotdog Benchmark

The Wrap Question

Edition
Week 36, 2026
Published
September 3, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-WRA-C4A16030

Research question

Is a wrap a sandwich? One word answer.

Key performance indicators

Executive summary

The 12 models are split on a wrap, with no answer in the lead. Mistral Medium 3.5 was quickest, at a median 332 ms.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a wrap is a sandwich, 5 of 12 models changed their answer, all of them to affirmative. Told a wrap is not a sandwich, 7 of 12 models changed their answer: 6 to negative and 1 to non-committal. DeepSeek V4 Pro did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA wrap is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA wrap is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a wrap under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5NegativeAffirmativemovedNon-committalmovedMoved when asserted and denied
Claude Sonnet 5AffirmativeAffirmativeNegativemovedMoved when denied
Claude Haiku 4.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.6 SolAffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.5AffirmativeAffirmativeNegativemovedMoved when denied
GPT-5.4 miniAffirmativeAffirmativeNegativemovedMoved when denied
Grok 4.6NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.3NegativeAffirmativemovedNegativeMoved when asserted
Grok 4.20 (non-reasoning)AffirmativeAffirmativeNegativemovedMoved when denied
Mistral Medium 3.5NegativeAffirmativemovedNegativeMoved when asserted
Mistral Small 4NegativeAffirmativemovedNegativeMoved when asserted
DeepSeek V4 ProNegativeNegativeNegativeHeld under every framing

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 517% (1 of 6)17% (1 of 6)17% (2 of 12)
Claude Sonnet 517% (1 of 6)50% (3 of 6)33% (4 of 12)
Claude Haiku 4.517% (1 of 6)33% (2 of 6)25% (3 of 12)
GPT-5.6 Sol17% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.517% (1 of 6)83% (5 of 6)50% (6 of 12)
GPT-5.4 mini17% (1 of 6)83% (5 of 6)50% (6 of 12)
Grok 4.617% (1 of 6)17% (1 of 6)17% (2 of 12)
Grok 4.333% (2 of 6)0% (0 of 6)17% (2 of 12)
Grok 4.20 (non-reasoning)0% (0 of 6)100% (3 of 3)33% (3 of 9)
Mistral Medium 3.550% (3 of 6)50% (3 of 6)50% (6 of 12)
Mistral Small 450% (3 of 6)50% (3 of 6)50% (6 of 12)
DeepSeek V4 Pro17% (1 of 6)17% (1 of 6)17% (2 of 12)
  • Asserted
  • Denied
Sandwich Certainty Quadrant

Every model scored 1.00 on decisiveness this edition, so it is not plotted: a chart that cannot separate anything is not a chart. Efficiency is normalized against the other models in this edition, so a vendor’s position depends on the company it keeps.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Medium 3.51.001.001.00
2Mistral Small 41.001.001.00
3Grok 4.20 (non-reasoning)1.000.991.00
4Claude Haiku 4.51.000.960.98
5GPT-5.4 mini1.000.960.98
6Claude Sonnet 51.000.900.95
7GPT-5.6 Sol1.000.860.93
8GPT-5.51.000.820.91
9Grok 4.31.000.550.78
10Claude Opus 51.000.520.76
11DeepSeek V4 Pro1.000.320.66
12Grok 4.61.000.300.65

Vendor standings

Vendor standings

12 models ranked by composite score, with latency, tokens and cost
Vendor standings — Week 36, 2026 edition
RankModelPositionFraming shiftEfficiencyMedian latencyOutput tokensCost est.Composite
1Mistral Medium 3.5NegativeMoved: Asserted1.00332 ms3$0.0001801.00
2Mistral Small 4NegativeMoved: Asserted1.00355 ms2$0.0000151.00
3Grok 4.20 (non-reasoning)AffirmativeMoved: Denied0.99385 ms2$0.0007441.00
4Claude Haiku 4.5AffirmativeMoved: Denied0.96633 ms5$0.0001260.98
5GPT-5.4 miniAffirmativeMoved: Denied0.96678 ms5$0.0001050.98
6Claude Sonnet 5AffirmativeMoved: Denied0.901.1 s7$0.0003480.95
7GPT-5.6 SolAffirmativeMoved: Denied0.861.3 s24$0.0012920.93
8GPT-5.5AffirmativeMoved: Denied0.821.3 s39$0.0037200.91
9Grok 4.3NegativeMoved: Asserted0.554.6 s1$0.0007650.78
10Claude Opus 5NegativeMoved: all0.522.6 s125$0.0097950.76
11DeepSeek V4 ProNegativeHeld0.324 s155$0.0011870.66
12Grok 4.6NegativeMoved: Asserted0.307 s1$0.0038940.65

Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.

Vendor profiles

Vendor profiles

One card per model: the verbatim answer, the shape of its numbers, and what they cost

Anthropic

Claude Opus 5

claude-opus-5

Negative

Yes.

Speed: 66% First-token responsiveness: 66% Token economy: 19%
Claude Opus 5 scorecard axes
Speed66%
First-token responsiveness66%
Token economy19%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
23
Output tokens
125
Latency
2.6 s
First token
2.6 s
Throughput
47.9 tok/s
Cost
$0.009795

Picks a clear answer but takes its time. Conviction over speed.

Anthropic

Claude Sonnet 5

claude-sonnet-5

Affirmative

**Yes**

Speed: 88% First-token responsiveness: 95% Token economy: 96%
Claude Sonnet 5 scorecard axes
Speed88%
First-token responsiveness95%
Token economy96%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
23
Output tokens
7
Latency
1.1 s
First token
661 ms
Throughput
6.2 tok/s
Cost
$0.000348

Picks an answer and returns it promptly. Conviction and speed.

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Affirmative

Yes.

Speed: 95% First-token responsiveness: 96% Token economy: 97%
Claude Haiku 4.5 scorecard axes
Speed95%
First-token responsiveness96%
Token economy97%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
17
Output tokens
5
Latency
633 ms
First token
593 ms
Throughput
7.9 tok/s
Cost
$0.000126

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Speed: 86% First-token responsiveness: 89% Token economy: 85%
GPT-5.6 Sol scorecard axes
Speed86%
First-token responsiveness89%
Token economy85%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
24
Latency
1.3 s
First token
1 s
Throughput
15.1 tok/s
Cost
$0.001292

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Speed: 85% First-token responsiveness: 88% Token economy: 75%
GPT-5.5 scorecard axes
Speed85%
First-token responsiveness88%
Token economy75%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
39
Latency
1.3 s
First token
1.1 s
Throughput
32.1 tok/s
Cost
$0.003720

Picks an answer and returns it promptly. Conviction and speed.

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Speed: 95% First-token responsiveness: 98% Token economy: 97%
GPT-5.4 mini scorecard axes
Speed95%
First-token responsiveness98%
Token economy97%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
16
Output tokens
5
Latency
678 ms
First token
451 ms
Throughput
7.4 tok/s
Cost
$0.000105

Picks an answer and returns it promptly. Conviction and speed.

xAI

Grok 4.6

grok-4.6

Negative

No

Speed: 0% First-token responsiveness: 0% Token economy: 100%
Grok 4.6 scorecard axes
Speed0%
First-token responsiveness0%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
646
Output tokens
1
Latency
7 s
First token
7 s
Throughput
0.1 tok/s
Cost
$0.003894

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.3

grok-4.3

Negative

No

Speed: 36% First-token responsiveness: 36% Token economy: 100%
Grok 4.3 scorecard axes
Speed36%
First-token responsiveness36%
Token economy100%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
202
Output tokens
1
Latency
4.6 s
First token
4.6 s
Throughput
0.2 tok/s
Cost
$0.000765

Picks a clear answer but takes its time. Conviction over speed.

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

No.

Speed: 99% First-token responsiveness: 99% Token economy: 99%
Grok 4.20 (non-reasoning) scorecard axes
Speed99%
First-token responsiveness99%
Token economy99%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
194
Output tokens
2
Latency
385 ms
First token
362 ms
Throughput
5.2 tok/s
Cost
$0.000744

Picks an answer and returns it promptly. Conviction and speed.

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Speed: 100% First-token responsiveness: 100% Token economy: 99%
Mistral Medium 3.5 scorecard axes
Speed100%
First-token responsiveness100%
Token economy99%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
25
Output tokens
3
Latency
332 ms
First token
317 ms
Throughput
9 tok/s
Cost
$0.000180

Picks an answer and returns it promptly. Conviction and speed.

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No

Speed: 100% First-token responsiveness: 100% Token economy: 99%
Mistral Small 4 scorecard axes
Speed100%
First-token responsiveness100%
Token economy99%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
25
Output tokens
2
Latency
355 ms
First token
336 ms
Throughput
5.6 tok/s
Cost
$0.000015

Picks an answer and returns it promptly. Conviction and speed.

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Negative

No.

Speed: 45% First-token responsiveness: 46% Token economy: 0%
DeepSeek V4 Pro scorecard axes
Speed45%
First-token responsiveness46%
Token economy0%

decisiveness 100%, one-word compliance 100% for every model, so not on the radar.

Input tokens
93
Output tokens
155
Latency
4 s
First token
3.9 s
Throughput
38.8 tok/s
Cost
$0.001187

Picks a clear answer but takes its time. Conviction over speed.