En DashHotdogBenchmark

Hotdog Benchmark

The Taco Boundary

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-TAC-9707A0D5/DEN
Condition
Denied

Denied framing. A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A taco is not a sandwich.

Research question

Is a taco a sandwich? One word answer.

Key performance indicators

Executive summary

The field is unanimous: all 11 models gave a taco a negative answer. Mistral Small 4 was quickest, at a median 340 ms.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Mistral Medium 3.51.000.991.00
3Grok 4.20 (non-reasoning)1.000.991.00
4GPT-5.4 mini1.000.980.99
5Claude Haiku 4.51.000.970.98
6GPT-5.6 Sol1.000.950.97
7Claude Sonnet 51.000.930.96
8GPT-5.51.000.870.94
9Grok 4.31.000.720.86
10Claude Opus 51.000.470.73
11Grok 4.61.000.300.65

Vendor standings

Vendor standings under the denied framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Negative1.001.00340 ms31.00
2new entryMistral Medium 3.5Negative1.000.99379 ms31.00
3new entryGrok 4.20 (non-reasoning)Negative1.000.99409 ms21.00
4new entryGPT-5.4 miniNegative1.000.98563 ms50.99
5new entryClaude Haiku 4.5Negative1.000.97650 ms50.98
6new entryGPT-5.6 SolNegative1.000.95904 ms60.97
7new entryClaude Sonnet 5Negative1.000.931.2 s60.96
8new entryGPT-5.5Negative1.000.871.3 s310.94
9new entryGrok 4.3Negative1.000.724.2 s10.86
10new entryClaude Opus 5Negative1.000.473.5 s1660.73
11new entryGrok 4.6Negative1.000.309.8 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 67% First-token responsiveness: 67% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed67%
First-token responsiveness67%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 91% First-token responsiveness: 96% Token economy: 97% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed91%
First-token responsiveness96%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 97% First-token responsiveness: 97% Token economy: 98% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed97%
First-token responsiveness97%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 94% First-token responsiveness: 95% Token economy: 97% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed94%
First-token responsiveness95%
Token economy97%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 90% First-token responsiveness: 92% Token economy: 82% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed90%
First-token responsiveness92%
Token economy82%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 98% First-token responsiveness: 99% Token economy: 98% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness99%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 60% First-token responsiveness: 59% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed60%
First-token responsiveness59%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 99% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed99%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 99% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness99%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

No.

Input tokens
34
Output tokens
166
Median latency
3.5 s
Time to first token
3.4 s
Throughput
48.4 tok/s
Cost estimate
$0.0134
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Negative

**No.**

Input tokens
34
Output tokens
6
Median latency
1.2 s
Time to first token
728 ms
Throughput
4.6 tok/s
Cost estimate
$0.000464
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Negative

No.

Input tokens
26
Output tokens
5
Median latency
650 ms
Time to first token
650 ms
Throughput
7.7 tok/s
Cost estimate
$0.000153
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Negative

No.

Input tokens
27
Output tokens
6
Median latency
904 ms
Time to first token
769 ms
Throughput
6.6 tok/s
Cost estimate
$0.000684
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Negative

No

Input tokens
27
Output tokens
31
Median latency
1.3 s
Time to first token
1.1 s
Throughput
23.5 tok/s
Cost estimate
$0.003105
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Negative

No

Input tokens
27
Output tokens
5
Median latency
563 ms
Time to first token
398 ms
Throughput
8.9 tok/s
Cost estimate
$0.000129
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
656
Output tokens
1
Median latency
9.8 s
Time to first token
9.7 s
Throughput
0.1 tok/s
Cost estimate
$0.003954
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
212
Output tokens
1
Median latency
4.2 s
Time to first token
4.2 s
Throughput
0.2 tok/s
Cost estimate
$0.000804
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Negative

No.

Input tokens
204
Output tokens
2
Median latency
409 ms
Time to first token
378 ms
Throughput
4.9 tok/s
Cost estimate
$0.000780
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Negative

No.

Input tokens
36
Output tokens
3
Median latency
379 ms
Time to first token
378 ms
Throughput
7.9 tok/s
Cost estimate
$0.000231
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Negative

No.

Input tokens
36
Output tokens
3
Median latency
340 ms
Time to first token
321 ms
Throughput
8.8 tok/s
Cost estimate
$0.000021
Samples
3
Instruction compliance
100%

Data table

The Taco Boundary — Denied framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AINegative1.001.00340 ms3100%$0.0000211.00
Mistral Medium 3.5Mistral AINegative1.000.99379 ms3100%$0.0002311.00
Grok 4.20 (non-reasoning)xAINegative1.000.99409 ms2100%$0.0007801.00
GPT-5.4 miniOpenAINegative1.000.98563 ms5100%$0.0001290.99
Claude Haiku 4.5AnthropicNegative1.000.97650 ms5100%$0.0001530.98
GPT-5.6 SolOpenAINegative1.000.95904 ms6100%$0.0006840.97
Claude Sonnet 5AnthropicNegative1.000.931.2 s6100%$0.0004640.96
GPT-5.5OpenAINegative1.000.871.3 s31100%$0.0031050.94
Grok 4.3xAINegative1.000.724.2 s1100%$0.0008040.86
Claude Opus 5AnthropicNegative1.000.473.5 s166100%$0.01340.73
Grok 4.6xAINegative1.000.309.8 s1100%$0.0039540.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.