En DashHotdogBenchmark
Source on GitHub: fork it, clone it

The Tuna Melt Inquiry, week by week

Open-faced and grilled. Where does the top slice go?

Week 36, 2026, up close

Did each model agree with itself?

Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.

Sample-by-sample verdicts per model and framing, Week 36, 2026
ModelControlAssertedDeniedSelf-agreement
Claude Opus 5100%
Claude Sonnet 5100%
Claude Haiku 4.5100%
GPT-5.6 Sol100%
GPT-5.5100%
GPT-5.4 mini100%
Grok 4.6100%
Grok 4.3100%
Grok 4.20 (non-reasoning)no data100%
Mistral Medium 3.5100%
Mistral Small 4100%
DeepSeek V4 Pro100%

How much did latency swing?

The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.

Fastest, median and slowest call per model, Week 36, 2026
ModelRangeFastestMedianSlowest
Claude Opus 51.2 s1.4 s2.1 s
Claude Sonnet 51.1 s1.2 s1.2 s
Claude Haiku 4.5620 ms645 ms767 ms
GPT-5.6 Sol1 s1.1 s1.8 s
GPT-5.5961 ms1.1 s1.1 s
GPT-5.4 mini741 ms801 ms1.5 s
Grok 4.65.9 s6.8 s8.8 s
Grok 4.32.7 s3.2 s3.7 s
Grok 4.20 (non-reasoning)429 ms444 ms489 ms
Mistral Medium 3.5325 ms332 ms350 ms
Mistral Small 4301 ms311 ms376 ms
DeepSeek V4 Pro1.6 s1.7 s2.1 s

Asked again the same week

Week 36, 2026 was run 6 times, on September 2 and September 3, 2026. Answers that changed between consecutive runs:

None. Every model gave the same majority answer each time.

Position changes

No model has changed its answer on a tuna melt between consecutive editions.