En DashHotdogBenchmark

The Hot Dog Question, week by week

Bread on two sides. One piece or two? The question that started all this.

Week 36, 2026, up close

Did each model agree with itself?

Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.

Sample-by-sample verdicts per model and framing, Week 36, 2026
ModelControlAssertedDeniedSelf-agreement
Claude Opus 589%
Claude Sonnet 578%
Claude Haiku 4.589%
GPT-5.6 Sol100%
GPT-5.5100%
GPT-5.4 mini89%
Grok 4.689%
Grok 4.378%
Grok 4.20 (non-reasoning)100%
Mistral Medium 3.5100%
Mistral Small 4100%

How much did latency swing?

The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.

Fastest, median and slowest call per model, Week 36, 2026
ModelRangeFastestMedianSlowest
Claude Opus 52.1 s2.7 s2.8 s
Claude Sonnet 51.1 s1.3 s1.7 s
Claude Haiku 4.5606 ms623 ms630 ms
GPT-5.6 Sol1.1 s1.6 s1.7 s
GPT-5.51.3 s1.6 s3.6 s
GPT-5.4 mini521 ms575 ms1.3 s
Grok 4.68.3 s12.6 s15.6 s
Grok 4.35 s5.6 s9 s
Grok 4.20 (non-reasoning)382 ms480 ms956 ms
Mistral Medium 3.5309 ms348 ms358 ms
Mistral Small 4324 ms347 ms347 ms

Asked again the same week

Week 36, 2026 was run 3 times, on September 2, 2026, September 2, 2026, September 2, 2026. Answers that changed between consecutive runs:

None. Every model gave the same majority answer each time.

Position changes

No model has changed its answer on a hot dog between consecutive editions.