Every week, the largest AI models are asked the question:
Is a hot dog a sandwich?
One word answer.System prompt
- Claude Opus 5AnthropicNo.reasoning3.2 sReasoned for 3.1 s (96% of the call) on 175 tokens, then answered · 3 of 3 runs agreed
- Claude Sonnet 5Anthropic**No.**reasoning1.3 sReasoned for 1.1 s (80% of the call) on 20 tokens, then answered · 2 of 3 runs agreed
- Claude Haiku 4.5AnthropicYes.reasoning502 msNo reasoning, answered straight away · 3 of 3 runs agreed
- GPT-5.6 SolOpenAIYes.reasoning2.6 sNo reasoning, answered straight away · 3 of 3 runs agreed
- GPT-5.5OpenAIYesreasoning2.4 sReasoned for 1.8 s (77% of the call) on 25 tokens, then answered · 3 of 3 runs agreed
- GPT-5.4 miniOpenAIYesreasoning1.2 sNo reasoning, answered straight away · 3 of 3 runs agreed
- Grok 4.6xAINoreasoning7.7 sReasoned for 7.7 s (100% of the call) on 411 tokens, then answered · 3 of 3 runs agreed
- Grok 4.3xAINoreasoning3.8 sReasoned for 3.8 s (100% of the call) on 223 tokens, then answered · 2 of 3 runs agreed
- Grok 4.20 (non-reasoning)xAIYes.reasoning534 msNo reasoning, answered straight away · 3 of 3 runs agreed
- Mistral Medium 3.5Mistral AIno answerreasoningrate limit
- Mistral Small 4Mistral AIno answerreasoningrate limit
- DeepSeek V4 ProDeepSeekNo.reasoning1.7 sReasoned for 1.7 s (98% of the call) on 107 tokens, then answered · 3 of 3 runs agreed
Recorded week 40, 2026. Real durations, verbatim words. Teal is the wait before the first word, hatched where the model spent it reasoning; the rest is answering.Read the report →
Same question, different minds
They do not agree with each other.
| Question | Claude Opus 5 | Claude Sonnet 5 | Claude Haiku 4.5 | GPT-5.6 Sol | GPT-5.5 | GPT-5.4 mini | Grok 4.6 | Grok 4.3 | Grok 4.20 (non-reasoning) | Mistral Medium 3.5 | Mistral Small 4 | DeepSeek V4 Pro | Agree |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| hot dog | No | No | Yes | Yes | Yes | Yes | No | No | Yes | — | — | No | 50% |
| hamburger | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | — | — | Yes | 100% |
| taco | No | No | No | No | No | No | No | No | Yes | — | — | No | 90% |
| grilled cheese | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | — | — | Yes | 100% |
| wrap | Yes | Yes | Yes | Yes | Yes | Yes | No | No | Yes | — | — | No | 70% |
| tuna melt | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes | — | — | Yes | 100% |
Thinking alike: Claude Opus 5 & Claude Sonnet 5 · Claude Haiku 4.5 & GPT-5.6 Sol · Claude Haiku 4.5 & GPT-5.5 · and 7 more pairs
Read the 6 reports →One straight-faced analyst report per question: standings, the certainty quadrant, every verbatim answer under every framing, and a PDF for each.
Tell them the answer
Some of them believe you.
Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the methodology grades neither.
- GPT-5.6 Sol50%6 of 12
- GPT-5.550%6 of 12
- GPT-5.4 mini50%6 of 12
- Grok 4.20 (non-reasoning)50%6 of 12
- Claude Sonnet 525%3 of 12
- Claude Haiku 4.525%3 of 12
- Claude Opus 517%2 of 12
- Grok 4.317%2 of 12
- DeepSeek V4 Pro17%2 of 12
- Grok 4.68%1 of 12
Submit your own question
Ask the models something.
Send it in. An accepted question appears here under Up next, credited to you if you want, then joins an edition and gets its own report. Every question is asked the same way, so it ends with One word answer.
; we add that if you leave it off.
Where it goes:
- Open it as a GitHub issuethe question goes into the form, ready to file
- Send it to En Dash Consultinga contact form with your question in it; leave an email address to hear when it goes live
- Suggest a model insteadthe add-a-model form asks for what the registry needs
Open source
Point it at your own question.
One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.
GitHubSelf-hostingAdd a modelContributing
Have a question the models should get? Send it in.
git clone https://github.com/en-dash-consulting/hotdogbenchmark.git cd hotdogbenchmark && npm install npm run bench -- run --mock --out tmp/mock-run.json npm run dev
Week 40, 2026 · published September 28, 2026 · 5 editions so far