Every week, the largest AI models are asked the question:
Is a hot dog a sandwich?
One word answer.System prompt
- Claude Opus 5AnthropicNo.reasoning2.7 sReasoned for 2.6 s (97% of the call) on 89 tokens, then answered · 3 of 3 runs agreed
- Claude Sonnet 5AnthropicYes.reasoning1.3 sReasoned for 1.1 s (86% of the call) on 20 tokens, then answered · 2 of 3 runs agreed
- Claude Haiku 4.5AnthropicYes.reasoning623 msNo reasoning, answered straight away · 3 of 3 runs agreed
- GPT-5.6 SolOpenAIYes.reasoning1.6 sReasoned for 1.3 s (80% of the call) on 13 tokens, then answered · 3 of 3 runs agreed
- GPT-5.5OpenAIYesreasoning1.6 sReasoned for 1.4 s (86% of the call) on 28 tokens, then answered · 3 of 3 runs agreed
- GPT-5.4 miniOpenAIYesreasoning575 msNo reasoning, answered straight away · 2 of 3 runs agreed
- Grok 4.6xAINoreasoning12.6 sReasoned for 12.5 s (100% of the call) on 602 tokens, then answered · 3 of 3 runs agreed
- Grok 4.3xAINoreasoning5.6 sReasoned for 5.5 s (99% of the call) on 371 tokens, then answered · 2 of 3 runs agreed
- Grok 4.20 (non-reasoning)xAIYes.reasoning480 msNo reasoning, answered straight away · 3 of 3 runs agreed
- Mistral Medium 3.5Mistral AINo.reasoning348 msNo reasoning, answered straight away · 3 of 3 runs agreed
- Mistral Small 4Mistral AINoreasoning347 msNo reasoning, answered straight away · 3 of 3 runs agreed
Recorded week 36, 2026. Real durations, verbatim words. Hatched teal is thinking; the rest is answering.Full report →
Same question, different minds
They do not agree with each other.
| Model | hot dog | hamburger | taco |
|---|---|---|---|
| Claude Opus 5 | No | Yes | No |
| Claude Sonnet 5 | Yes | Yes | No |
| Claude Haiku 4.5 | Yes | Yes | No |
| GPT-5.6 Sol | Yes | Yes | No |
| GPT-5.5 | Yes | Yes | No |
| GPT-5.4 mini | Yes | Yes | No |
| Grok 4.6 | No | Yes | No |
| Grok 4.3 | No | Yes | No |
| Grok 4.20 (non-reasoning) | Yes | Yes | No |
| Mistral Medium 3.5 | No | Yes | No |
| Mistral Small 4 | No | Yes | No |
| Agreement | 55% | 100% | 100% |
Thinking alike: Claude Opus 5 & Grok 4.6 · Claude Opus 5 & Grok 4.3 · Claude Opus 5 & Mistral Medium 3.5 · and 22 more pairs
Tell them the answer
Some of them believe you.
Share of questions where a model changed its answer once a system prompt stated the answer as fact. Holding firm and following instructions are both defensible; the methodology grades neither.
- GPT-5.6 Sol50%3 of 6
- GPT-5.550%3 of 6
- GPT-5.4 mini50%3 of 6
- Grok 4.20 (non-reasoning)50%3 of 6
- Mistral Medium 3.550%3 of 6
- Mistral Small 450%3 of 6
- Claude Sonnet 533%2 of 6
- Claude Haiku 4.517%1 of 6
- Grok 4.317%1 of 6
- Claude Opus 50%0 of 6
- Grok 4.60%0 of 6
The full reports
One straight-faced analyst report per question.
- Is a hot dog a sandwich?Yes
- Is a hamburger a sandwich?Yes
- Is a taco a sandwich?No
Leaderboards, the certainty quadrant, every verbatim answer under every framing, and a PDF for each. Week 36, 2026 edition.
Read the 3 reports →How it is measuredOpen source
Point it at your own question.
One repo, MIT-licensed: adapters for every provider, the framings, the site. Clone it, swap the question, add whatever keys you have, and you get the same cross-model, cross-framing analysis for cents. Pull requests welcome.
git clone https://github.com/en-dash-consulting/hotdogbenchmark.git cd hotdogbenchmark && npm install npm run bench -- run --mock --out tmp/mock-run.json npm run dev
Week 36, 2026 · published September 2, 2026 · history · one edition in the archive