Week 36, 2026 edition
An archived record of every question asked in this edition. Current reports are on the home page.
The Hot Dog Question
Key performance indicators
Models evaluated
11
Consensus position
Affirmative
55% of the field
Median latency
1.3 s
Median output tokens
5
Instruction compliance
97%
Executive summary
A affirmative answer on a hot dog has majority support this week: 6 of 11 models (55%). Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 disagree. Mistral Small 4 was quickest, at a median 347 ms. 97% of answers were actually one word, as asked.
Key findings
- Consensus. 55% of models (6 of 11) say affirmative.
- Dissent. Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 went the other way.
- Response latency. Mistral Small 4 answered in a median 347 ms; Grok 4.6 took 12.6 seconds. That is a 36.3× spread, mostly thinking time.
- Response length. Claude Opus 5 used the most output tokens, a median of 93, for a question that asked for one word.
- Instruction compliance. Claude Sonnet 5 did not always keep it to one word. Following the instruction and picking an answer are scored separately; they are different skills.
- Composite standing. Mistral Small 4 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
The Hamburger Control
Key performance indicators
Models evaluated
11
Consensus position
Affirmative
100% of the field
Median latency
1.1 s
Median output tokens
5
Instruction compliance
100%
Executive summary
The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.
Key findings
- Consensus. Everyone agrees: a hamburger gets a affirmative from every model, unanimous. That does not happen often.
- Response latency. Mistral Small 4 answered in a median 340 ms; Grok 4.6 took 7.0 seconds. That is a 20.8× spread, mostly thinking time.
- Response length. Claude Opus 5 used the most output tokens, a median of 123, for a question that asked for one word.
- Instruction compliance. Every model kept it to one word. Nice.
- Composite standing. Mistral Small 4 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
The Taco Boundary
Key performance indicators
Models evaluated
11
Consensus position
Negative
100% of the field
Median latency
940 ms
Median output tokens
3
Instruction compliance
100%
Executive summary
The field is unanimous: all 11 models gave a taco a negative answer. Mistral Medium 3.5 was quickest, at a median 365 ms.
Key findings
- Consensus. Everyone agrees: a taco gets a negative from every model, unanimous. That does not happen often.
- Response latency. Mistral Medium 3.5 answered in a median 365 ms; Grok 4.6 took 10.5 seconds. That is a 28.7× spread, mostly thinking time.
- Response length. Claude Opus 5 used the most output tokens, a median of 68, for a question that asked for one word.
- Instruction compliance. Every model kept it to one word. Nice.
- Composite standing. Mistral Medium 3.5 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.