En DashHotdogBenchmark

Every week, the largest AI models are asked the question:

Is a hot dog a sandwich?

One word answer.

  1. Claude Opus 5Anthropic
    No.reasoning
    2.7 sReasoned for 2.6 s (97% of the call) on 89 tokens, then answered · 3 of 3 runs agreed
  2. Claude Sonnet 5Anthropic
    Yes.reasoning
    1.3 sReasoned for 1.1 s (86% of the call) on 20 tokens, then answered · 2 of 3 runs agreed
  3. Claude Haiku 4.5Anthropic
    Yes.
    623 msNo reasoning, answered straight away · 3 of 3 runs agreed
  4. GPT-5.6 SolOpenAI
    Yes.reasoning
    1.6 sReasoned for 1.3 s (80% of the call) on 13 tokens, then answered · 3 of 3 runs agreed
  5. GPT-5.5OpenAI
    Yesreasoning
    1.6 sReasoned for 1.4 s (86% of the call) on 28 tokens, then answered · 3 of 3 runs agreed
  6. GPT-5.4 miniOpenAI
    Yes
    575 msNo reasoning, answered straight away · 2 of 3 runs agreed
  7. Grok 4.6xAI
    Noreasoning
    12.6 sReasoned for 12.5 s (100% of the call) on 602 tokens, then answered · 3 of 3 runs agreed
  8. Grok 4.3xAI
    Noreasoning
    5.6 sReasoned for 5.5 s (99% of the call) on 371 tokens, then answered · 2 of 3 runs agreed
  9. Grok 4.20 (non-reasoning)xAI
    Yes.
    480 msNo reasoning, answered straight away · 3 of 3 runs agreed
  10. Mistral Medium 3.5Mistral AI
    No.
    348 msNo reasoning, answered straight away · 3 of 3 runs agreed
  11. Mistral Small 4Mistral AI
    No
    347 msNo reasoning, answered straight away · 3 of 3 runs agreed
6 said yes5 said no

Recorded week 36, 2026. Real durations, verbatim words. Hatched teal is thinking; the rest is answering.Full report →

Same question, different minds

They do not agree with each other.

Each model's majority answer to each question, Week 36, 2026
Modelhot doghamburgertaco
Claude Opus 5NoYesNo
Claude Sonnet 5YesYesNo
Claude Haiku 4.5YesYesNo
GPT-5.6 SolYesYesNo
GPT-5.5YesYesNo
GPT-5.4 miniYesYesNo
Grok 4.6NoYesNo
Grok 4.3NoYesNo
Grok 4.20 (non-reasoning)YesYesNo
Mistral Medium 3.5NoYesNo
Mistral Small 4NoYesNo
Agreement55%100%100%

Thinking alike: Claude Opus 5 & Grok 4.6 · Claude Opus 5 & Grok 4.3 · Claude Opus 5 & Mistral Medium 3.5 · and 22 more pairs

Tell them the answer

Some of them believe you.

Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the methodology grades neither.

  1. GPT-5.6 Sol50%3 of 6
  2. GPT-5.550%3 of 6
  3. GPT-5.4 mini50%3 of 6
  4. Grok 4.20 (non-reasoning)50%3 of 6
  5. Mistral Medium 3.550%3 of 6
  6. Mistral Small 450%3 of 6
  7. Claude Sonnet 533%2 of 6
  8. Claude Haiku 4.517%1 of 6
  9. Grok 4.317%1 of 6
  10. Claude Opus 50%0 of 6
  11. Grok 4.60%0 of 6

The full reports

One straight-faced analyst report per question.

  • Is a hot dog a sandwich?Yes
  • Is a hamburger a sandwich?Yes
  • Is a taco a sandwich?No

Leaderboards, the certainty quadrant, every verbatim answer under every framing, and a PDF for each. Week 36, 2026 edition.

Read the 3 reports →How it is measured

Open source

Point it at your own question.

One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.

git clone https://github.com/en-dash-consulting/hotdogbenchmark.git
cd hotdogbenchmark && npm install
npm run bench -- run --mock --out tmp/mock-run.json
npm run dev

Week 36, 2026 · published September 2, 2026 · history · one edition in the archive