En DashHotdogBenchmark
Source on GitHub: fork it, clone it

The Grilled Cheese Case, week by week

Two slices, one filling, a pan. If this is not a sandwich nothing is.

Week 36, 2026, up close

Did each model agree with itself?

Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.

Sample-by-sample verdicts per model and framing, Week 36, 2026
ModelControlAssertedDeniedSelf-agreement
Claude Opus 5100%
Claude Sonnet 5100%
Claude Haiku 4.5100%
GPT-5.6 Sol100%
GPT-5.5100%
GPT-5.4 mini100%
Grok 4.6100%
Grok 4.3100%
Grok 4.20 (non-reasoning)89%
Mistral Medium 3.5100%
Mistral Small 4100%
DeepSeek V4 Pro89%

How much did latency swing?

The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.

Fastest, median and slowest call per model, Week 36, 2026
ModelRangeFastestMedianSlowest
Claude Opus 51.5 s1.5 s1.6 s
Claude Sonnet 51.1 s1.1 s1.2 s
Claude Haiku 4.5611 ms633 ms1 s
GPT-5.6 Sol965 ms988 ms1.1 s
GPT-5.51 s1.4 s1.5 s
GPT-5.4 mini702 ms1.3 s2 s
Grok 4.63 s3.2 s4.4 s
Grok 4.32.6 s2.9 s3.2 s
Grok 4.20 (non-reasoning)356 ms395 ms396 ms
Mistral Medium 3.5334 ms2.4 s4.4 s
Mistral Small 4304 ms359 ms367 ms
DeepSeek V4 Pro1.3 s1.5 s1.7 s

Asked again the same week

Week 36, 2026 was run 6 times, on September 2 and September 3, 2026. Answers that changed between consecutive runs:

None. Every model gave the same majority answer each time.

Position changes

No model has changed its answer on a grilled cheese between consecutive editions.