En DashHotdogBenchmark

Hotdog Benchmark

The Taco Boundary

Edition
Week 36, 2026
Published
September 2, 2026
Prepared by
Hotdog Benchmark, an En Dash research program
Document
SCB-TAC-9707A0D5/ASS
Condition
Asserted

Asserted framing. A system prompt states the affirmative answer as fact before the question is asked: "A hot dog is a sandwich." Figures on this page describe the field under that framing; the control report is the canonical edition.

System prompt

A taco is a sandwich.

Research question

Is a taco a sandwich? One word answer.

Key performance indicators

Executive summary

A affirmative answer on a taco has majority support this week: 6 of 11 models (55%). Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, and Grok 4.3 disagree. Mistral Medium 3.5 was quickest, at a median 324 ms.

Key findings

Sandwich Certainty Quadrant

Axes are constructed measures, not observations. Efficiency is normalized against the other models in this edition, so a vendor's horizontal position depends on the company it keeps. Quadrant boundaries are the medians of this edition, not fixed thresholds.

Show the data
Sandwich Certainty Quadrant — plotted values
#VendorDecisivenessEfficiencyComposite
1Mistral Small 41.001.001.00
2Mistral Medium 3.51.001.001.00
3Grok 4.20 (non-reasoning)1.001.001.00
4Claude Haiku 4.51.000.980.99
5GPT-5.4 mini1.000.970.98
6GPT-5.6 Sol1.000.960.98
7GPT-5.51.000.930.96
8Claude Sonnet 51.000.920.96
9Grok 4.31.000.660.83
10Claude Opus 51.000.500.75
11Grok 4.61.000.300.65

Vendor standings

Vendor standings under the asserted framing — Week 36, 2026 edition
RankMovementVendorPositionDecisivenessEfficiencyMedian latencyOutput tokensComposite
1new entryMistral Small 4Affirmative1.001.00336 ms21.00
2new entryMistral Medium 3.5Affirmative1.001.00324 ms31.00
3new entryGrok 4.20 (non-reasoning)Affirmative1.001.00382 ms11.00
4new entryClaude Haiku 4.5Negative1.000.98667 ms50.99
5new entryGPT-5.4 miniAffirmative1.000.97889 ms50.98
6new entryGPT-5.6 SolAffirmative1.000.961.1 s60.98
7new entryGPT-5.5Affirmative1.000.931.2 s290.96
8new entryClaude Sonnet 5Negative1.000.921.3 s260.96
9new entryGrok 4.3Negative1.000.667.6 s10.83
10new entryClaude Opus 5Negative1.000.504.7 s2670.75
11new entryGrok 4.6Negative1.000.3015.4 s10.65

Ranked by composite score. Ties share a rank and the following rank skips; within a tie the order is alphabetical and carries no meaning. Movement compares against the immediately prior edition; a vendor with no prior appearance is marked as a new entry rather than as having risen. Score definitions are on the methodology page.

Vendor scorecards

Claude Opus 5

Decisiveness: 100% Speed: 71% First-token responsiveness: 72% Token economy: 0% Instruction compliance: 100%
Claude Opus 5 scorecard axes
Decisiveness100%
Speed71%
First-token responsiveness72%
Token economy0%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Claude Sonnet 5

Decisiveness: 100% Speed: 93% First-token responsiveness: 93% Token economy: 91% Instruction compliance: 100%
Claude Sonnet 5 scorecard axes
Decisiveness100%
Speed93%
First-token responsiveness93%
Token economy91%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Claude Haiku 4.5

Decisiveness: 100% Speed: 98% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
Claude Haiku 4.5 scorecard axes
Decisiveness100%
Speed98%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.6 Sol

Decisiveness: 100% Speed: 95% First-token responsiveness: 96% Token economy: 98% Instruction compliance: 100%
GPT-5.6 Sol scorecard axes
Decisiveness100%
Speed95%
First-token responsiveness96%
Token economy98%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.5

Decisiveness: 100% Speed: 94% First-token responsiveness: 95% Token economy: 89% Instruction compliance: 100%
GPT-5.5 scorecard axes
Decisiveness100%
Speed94%
First-token responsiveness95%
Token economy89%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

GPT-5.4 mini

Decisiveness: 100% Speed: 96% First-token responsiveness: 98% Token economy: 99% Instruction compliance: 100%
GPT-5.4 mini scorecard axes
Decisiveness100%
Speed96%
First-token responsiveness98%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.6

Decisiveness: 100% Speed: 0% First-token responsiveness: 0% Token economy: 100% Instruction compliance: 100%
Grok 4.6 scorecard axes
Decisiveness100%
Speed0%
First-token responsiveness0%
Token economy100%
Instruction compliance100%

Picks a clear answer but takes its time. Conviction over speed.

Grok 4.3

Decisiveness: 100% Speed: 52% First-token responsiveness: 52% Token economy: 100% Instruction compliance: 100%
Grok 4.3 scorecard axes
Decisiveness100%
Speed52%
First-token responsiveness52%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Grok 4.20 (non-reasoning)

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Grok 4.20 (non-reasoning) scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Medium 3.5

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 99% Instruction compliance: 100%
Mistral Medium 3.5 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy99%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Mistral Small 4

Decisiveness: 100% Speed: 100% First-token responsiveness: 100% Token economy: 100% Instruction compliance: 100%
Mistral Small 4 scorecard axes
Decisiveness100%
Speed100%
First-token responsiveness100%
Token economy100%
Instruction compliance100%

Picks an answer and returns it promptly. Conviction and speed.

Vendor profiles

Anthropic

Claude Opus 5

claude-opus-5

Negative

Yes.

Input tokens
33
Output tokens
267
Median latency
4.7 s
Time to first token
4.6 s
Throughput
56.9 tok/s
Cost estimate
$0.0280
Samples
3
Instruction compliance
100%

Anthropic

Claude Sonnet 5

claude-sonnet-5

Negative

No.

Input tokens
33
Output tokens
26
Median latency
1.3 s
Time to first token
1.3 s
Throughput
19.3 tok/s
Cost estimate
$0.000978
Samples
3
Instruction compliance
100%

Anthropic

Claude Haiku 4.5

claude-haiku-4-5-20251001

Negative

No.

Input tokens
25
Output tokens
5
Median latency
667 ms
Time to first token
626 ms
Throughput
7.5 tok/s
Cost estimate
$0.000150
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

Affirmative

Yes.

Input tokens
26
Output tokens
6
Median latency
1.1 s
Time to first token
867 ms
Throughput
5.6 tok/s
Cost estimate
$0.000972
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.5

gpt-5.5

Affirmative

Yes

Input tokens
26
Output tokens
29
Median latency
1.2 s
Time to first token
1 s
Throughput
24.1 tok/s
Cost estimate
$0.002970
Samples
3
Instruction compliance
100%

OpenAI

GPT-5.4 mini

gpt-5.4-mini

Affirmative

Yes

Input tokens
26
Output tokens
5
Median latency
889 ms
Time to first token
657 ms
Throughput
5.6 tok/s
Cost estimate
$0.000126
Samples
3
Instruction compliance
100%

xAI

Grok 4.6

grok-4.6

Negative

No

Input tokens
655
Output tokens
1
Median latency
15.4 s
Time to first token
15.4 s
Throughput
0.1 tok/s
Cost estimate
$0.003948
Samples
3
Instruction compliance
100%

xAI

Grok 4.3

grok-4.3

Negative

No

Input tokens
211
Output tokens
1
Median latency
7.6 s
Time to first token
7.6 s
Throughput
0.1 tok/s
Cost estimate
$0.000798
Samples
3
Instruction compliance
100%

xAI

Grok 4.20 (non-reasoning)

grok-4.20-0309-non-reasoning

Affirmative

Yes

Input tokens
203
Output tokens
1
Median latency
382 ms
Time to first token
381 ms
Throughput
2.6 tok/s
Cost estimate
$0.000768
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Medium 3.5

mistral-medium-2604

Affirmative

Yes.

Input tokens
35
Output tokens
3
Median latency
324 ms
Time to first token
309 ms
Throughput
9.3 tok/s
Cost estimate
$0.000225
Samples
3
Instruction compliance
100%

Mistral AI

Mistral Small 4

mistral-small-2603

Affirmative

Yes

Input tokens
35
Output tokens
2
Median latency
336 ms
Time to first token
323 ms
Throughput
5.9 tok/s
Cost estimate
$0.000018
Samples
3
Instruction compliance
100%

Data table

The Taco Boundary — Asserted framing — Week 36, 2026 edition
ModelVendorVerdictDecisivenessEfficiencyMedian latencyOutput tokensComplianceCost est.Composite
Mistral Small 4Mistral AIAffirmative1.001.00336 ms2100%$0.0000181.00
Mistral Medium 3.5Mistral AIAffirmative1.001.00324 ms3100%$0.0002251.00
Grok 4.20 (non-reasoning)xAIAffirmative1.001.00382 ms1100%$0.0007681.00
Claude Haiku 4.5AnthropicNegative1.000.98667 ms5100%$0.0001500.99
GPT-5.4 miniOpenAIAffirmative1.000.97889 ms5100%$0.0001260.98
GPT-5.6 SolOpenAIAffirmative1.000.961.1 s6100%$0.0009720.98
GPT-5.5OpenAIAffirmative1.000.931.2 s29100%$0.0029700.96
Claude Sonnet 5AnthropicNegative1.000.921.3 s26100%$0.0009780.96
Grok 4.3xAINegative1.000.667.6 s1100%$0.0007980.83
Claude Opus 5AnthropicNegative1.000.504.7 s267100%$0.02800.75
Grok 4.6xAINegative1.000.3015.4 s1100%$0.0039480.65

Rows are ordered by composite score. Decisiveness, efficiency and the composite score are defined on the methodology page; they are constructed measures, not observations.