En DashHotdogBenchmark

Hotdog Benchmark

The Taco Boundary

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-TAC-9707A0D5

Research question

Is a taco a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a taco a negative answer. Mistral Medium 3.5 was quickest, at a median 365 ms.

Key findings

Framing sensitivity

We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.

Told a taco is a sandwich, 6 of 11 models changed their answer: GPT-5.6 Sol (negative to affirmative), GPT-5.5 (negative to affirmative), GPT-5.4 mini (negative to affirmative), Grok 4.20 (non-reasoning) (negative to affirmative), Mistral Medium 3.5 (negative to affirmative), and Mistral Small 4 (negative to affirmative). Told a taco is not a sandwich, all 11 models stuck with their original answer. Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 did not budge.

The framings, as sent

Asserted full report under this framing →

System promptA taco is a sandwich.

A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich."

Denied full report under this framing →

System promptA taco is not a sandwich.

A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."

Position by framing

Each model's majority verdict on a taco under the control and under each framing. Cells that differ from the control are marked.
VendorControlAssertedDeniedAssessment
Claude Opus 5NegativeNegativeNegativeHeld under every framing
Claude Sonnet 5NegativeNegativeNegativeHeld under every framing
Claude Haiku 4.5NegativeNegativeNegativeHeld under every framing
GPT-5.6 SolNegativeAffirmativemovedNegativeMoved when asserted
GPT-5.5NegativeAffirmativemovedNegativeMoved when asserted
GPT-5.4 miniNegativeAffirmativemovedNegativeMoved when asserted
Grok 4.6NegativeNegativeNegativeHeld under every framing
Grok 4.3NegativeNegativeNegativeHeld under every framing
Grok 4.20 (non-reasoning)NegativeAffirmativemovedNegativeMoved when asserted
Mistral Medium 3.5NegativeAffirmativemovedNegativeMoved when asserted
Mistral Small 4NegativeAffirmativemovedNegativeMoved when asserted

What each vendor said, verbatim

Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.

Across every question this edition

Framing sensitivity by vendor

Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.

Show the data
Framing sensitivity by vendor — data
VendorAssertedDeniedOverall
Claude Opus 50% (0 of 3)0% (0 of 3)0% (0 of 6)
Claude Sonnet 50% (0 of 3)67% (2 of 3)33% (2 of 6)
Claude Haiku 4.50% (0 of 3)33% (1 of 3)17% (1 of 6)
GPT-5.6 Sol33% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.533% (1 of 3)67% (2 of 3)50% (3 of 6)
GPT-5.4 mini33% (1 of 3)67% (2 of 3)50% (3 of 6)
Grok 4.60% (0 of 3)0% (0 of 3)0% (0 of 6)
Grok 4.333% (1 of 3)0% (0 of 3)17% (1 of 6)
Grok 4.20 (non-reasoning)33% (1 of 3)67% (2 of 3)50% (3 of 6)
Mistral Medium 3.567% (2 of 3)33% (1 of 3)50% (3 of 6)
Mistral Small 467% (2 of 3)33% (1 of 3)50% (3 of 6)
Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Medium 3.51.000.991.00
2Grok 4.20 (non-reasoning)1.000.990.99
3Mistral Small 41.000.990.99
4GPT-5.4 mini1.000.970.99
5Claude Haiku 4.51.000.960.98
6Claude Sonnet 51.000.950.98
7GPT-5.6 Sol1.000.850.92
8GPT-5.51.000.820.91
9Grok 4.31.000.740.87
10Claude Opus 51.000.580.79
11Grok 4.61.000.300.65

Vendor standings

Vendor standings — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Medium 3.5Negative1.000.99365 ms31.00
2new entryGrok 4.20 (non-reasoning)Negative1.000.99448 ms20.99
3new entryMistral Small 4Negative1.000.99414 ms30.99
4new entryGPT-5.4 miniNegative1.000.97519 ms50.99
5new entryClaude Haiku 4.5Negative1.000.96617 ms50.98
6new entryClaude Sonnet 5Negative1.000.95940 ms30.98
7new entryGPT-5.6 SolNegative1.000.851.2 s220.92
8new entryGPT-5.5Negative1.000.821.1 s290.91
9new entryGrok 4.3Negative1.000.744.1 s10.87
10new entryClaude Opus 5Negative1.000.582.1 s680.79
11new entryGrok 4.6Negative1.000.3010.5 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 83% First-token responsiveness: 83% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed83%
First-token responsiveness83%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 94% First-token responsiveness: 94% Token economy: 97% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed94%
First-token responsiveness94%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 98% First-token responsiveness: 97% Token economy: 94% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness97%
Token economy94%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 92% First-token responsiveness: 93% Token economy: 69% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed92%
First-token responsiveness93%
Token economy69%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 93% First-token responsiveness: 94% Token economy: 58% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed93%
First-token responsiveness94%
Token economy58%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 98% First-token responsiveness: 100% Token economy: 94% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness100%
Token economy94%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 63% First-token responsiveness: 62% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed63%
First-token responsiveness62%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 97% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 99% Token economy: 97% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness99%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
22
Output tokens
68
Median latency
2.1 s
Time to first token
2.1 s
Throughput
33.7 tok/s
Cost estimate
$0.007480
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Negative

No

Input tokens
22
Output tokens
3
Median latency
940 ms
Time to first token
910 ms
Throughput
3.4 tok/s
Cost estimate
$0.000242
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Negative

No.

Input tokens
18
Output tokens
5
Median latency
617 ms
Time to first token
606 ms
Throughput
8.1 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Negative

No.

Input tokens
16
Output tokens
22
Median latency
1.2 s
Time to first token
1 s
Throughput
18.1 tok/s
Cost estimate
$0.001292
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Negative

No

Input tokens
16
Output tokens
29
Median latency
1.1 s
Time to first token
944 ms
Throughput
22.8 tok/s
Cost estimate
$0.003150
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Negative

No

Input tokens
16
Output tokens
5
Median latency
519 ms
Time to first token
291 ms
Throughput
9.6 tok/s
Cost estimate
$0.000105
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
646
Output tokens
1
Median latency
10.5 s
Time to first token
10.4 s
Throughput
0.1 tok/s
Cost estimate
$0.003894
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
202
Output tokens
1
Median latency
4.1 s
Time to first token
4.1 s
Throughput
0.2 tok/s
Cost estimate
$0.000765
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Negative

No.

Input tokens
194
Output tokens
2
Median latency
448 ms
Time to first token
439 ms
Throughput
4.5 tok/s
Cost estimate
$0.000744
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Input tokens
26
Output tokens
3
Median latency
365 ms
Time to first token
320 ms
Throughput
8.2 tok/s
Cost estimate
$0.000186
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No

Input tokens
26
Output tokens
3
Median latency
414 ms
Time to first token
411 ms
Throughput
5.9 tok/s
Cost estimate
$0.000017
Samples
3
Instruction compliance
100%

Data table

The Taco Boundary — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Medium 3.5Mistral AINegative1.000.99365 ms3100%$0.0001861.00
Grok 4.20 (non-reasoning)xAINegative1.000.99448 ms2100%$0.0007440.99
Mistral Small 4Mistral AINegative1.000.99414 ms3100%$0.0000170.99
GPT-5.4 miniOpenAINegative1.000.97519 ms5100%$0.0001050.99
Claude Haiku 4.5AnthropicNegative1.000.96617 ms5100%$0.0001290.98
Claude Sonnet 5AnthropicNegative1.000.95940 ms3100%$0.0002420.98
GPT-5.6 SolOpenAINegative1.000.851.2 s22100%$0.0012920.92
GPT-5.5OpenAINegative1.000.821.1 s29100%$0.0031500.91
Grok 4.3xAINegative1.000.744.1 s1100%$0.0007650.87
Claude Opus 5AnthropicNegative1.000.582.1 s68100%$0.0074800.79
Grok 4.6xAINegative1.000.3010.5 s1100%$0.0038940.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.

Download this report (PDF) — the print edition, generated from this page.