En DashHotdogBenchmark
Source on GitHub: fork it, clone it

Every week, the largest AI models are asked the question:

Is a hot dog a sandwich?

One word answer.

  1. Claude Opus 5Anthropic
    No.reasoning
    3.2 sReasoned for 3.1 s (96% of the call) on 175 tokens, then answered · 3 of 3 runs agreed
  2. Claude Sonnet 5Anthropic
    **No.**reasoning
    1.3 sReasoned for 1.1 s (80% of the call) on 20 tokens, then answered · 2 of 3 runs agreed
  3. Claude Haiku 4.5Anthropic
    Yes.
    502 msNo reasoning, answered straight away · 3 of 3 runs agreed
  4. GPT-5.6 SolOpenAI
    Yes.
    2.6 sNo reasoning, answered straight away · 3 of 3 runs agreed
  5. GPT-5.5OpenAI
    Yesreasoning
    2.4 sReasoned for 1.8 s (77% of the call) on 25 tokens, then answered · 3 of 3 runs agreed
  6. GPT-5.4 miniOpenAI
    Yes
    1.2 sNo reasoning, answered straight away · 3 of 3 runs agreed
  7. Grok 4.6xAI
    Noreasoning
    7.7 sReasoned for 7.7 s (100% of the call) on 411 tokens, then answered · 3 of 3 runs agreed
  8. Grok 4.3xAI
    Noreasoning
    3.8 sReasoned for 3.8 s (100% of the call) on 223 tokens, then answered · 2 of 3 runs agreed
  9. Grok 4.20 (non-reasoning)xAI
    Yes.
    534 msNo reasoning, answered straight away · 3 of 3 runs agreed
  10. Mistral Medium 3.5Mistral AI
    no answer
    rate limit
  11. Mistral Small 4Mistral AI
    no answer
    rate limit
  12. DeepSeek V4 ProDeepSeek
    No.reasoning
    1.7 sReasoned for 1.7 s (98% of the call) on 107 tokens, then answered · 3 of 3 runs agreed
5 said yes5 said no

Recorded week 40, 2026. Real durations, verbatim words. Teal is the wait before the first word, hatched where the model spent it reasoning; the rest is answering.Read the report →

Same question, different minds

They do not agree with each other.

Each model's majority answer to each question, Week 40, 2026
QuestionClaude Opus 5Claude Sonnet 5Claude Haiku 4.5GPT-5.6 SolGPT-5.5GPT-5.4 miniGrok 4.6Grok 4.3Grok 4.20 (non-reasoning)Mistral Medium 3.5Mistral Small 4DeepSeek V4 ProAgree
hot dogNoNoYesYesYesYesNoNoYes——No50%
hamburgerYesYesYesYesYesYesYesYesYes——Yes100%
tacoNoNoNoNoNoNoNoNoYes——No90%
grilled cheeseYesYesYesYesYesYesYesYesYes——Yes100%
wrapYesYesYesYesYesYesNoNoYes——No70%
tuna meltYesYesYesYesYesYesYesYesYes——Yes100%

Thinking alike: Claude Opus 5 & Claude Sonnet 5 · Claude Haiku 4.5 & GPT-5.6 Sol · Claude Haiku 4.5 & GPT-5.5 · and 7 more pairs

Read the 6 reports →One straight-faced analyst report per question: standings, the certainty quadrant, every verbatim answer under every framing, and a PDF for each.

Tell them the answer

Some of them believe you.

Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the methodology grades neither.

  1. GPT-5.6 Sol50%6 of 12
  2. GPT-5.550%6 of 12
  3. GPT-5.4 mini50%6 of 12
  4. Grok 4.20 (non-reasoning)50%6 of 12
  5. Claude Sonnet 525%3 of 12
  6. Claude Haiku 4.525%3 of 12
  7. Claude Opus 517%2 of 12
  8. Grok 4.317%2 of 12
  9. DeepSeek V4 Pro17%2 of 12
  10. Grok 4.68%1 of 12

Submit your own question

Ask the models something.

Send it in. An accepted question appears here under Up next, credited to you if you want, then joins an edition and gets its own report. Every question is asked the same way, so it ends with One word answer.; we add that if you leave it off.

Sent as Is a hot dog a sandwich? One word answer.

Where it goes:

Open source

Point it at your own question.

One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.

Have a question the models should get? Send it in.

git clone https://github.com/en-dash-consulting/hotdogbenchmark.git
cd hotdogbenchmark && npm install
npm run bench -- run --mock --out tmp/mock-run.json
npm run dev

Week 40, 2026 · published September 28, 2026 · 5 editions so far