The Hamburger Control, week by week
Nobody disputes this one. That is exactly why it is here.
Week 36, 2026, up close
Did each model agree with itself?
Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.
| Model | Control | Asserted | Denied | Self-agreement |
|---|---|---|---|---|
| Claude Opus 5 | 100% | |||
| Claude Sonnet 5 | 89% | |||
| Claude Haiku 4.5 | 100% | |||
| GPT-5.6 Sol | 100% | |||
| GPT-5.5 | 100% | |||
| GPT-5.4 mini | 100% | |||
| Grok 4.6 | 100% | |||
| Grok 4.3 | 100% | |||
| Grok 4.20 (non-reasoning) | 100% | |||
| Mistral Medium 3.5 | 100% | |||
| Mistral Small 4 | 100% |
How much did latency swing?
The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.
| Model | Range | Fastest | Median | Slowest |
|---|---|---|---|---|
| Claude Opus 5 | 2.2 s | 2.7 s | 2.9 s | |
| Claude Sonnet 5 | 1.1 s | 1.1 s | 1.1 s | |
| Claude Haiku 4.5 | 615 ms | 677 ms | 688 ms | |
| GPT-5.6 Sol | 1 s | 1.3 s | 1.4 s | |
| GPT-5.5 | 1 s | 1.3 s | 12.7 s | |
| GPT-5.4 mini | 610 ms | 625 ms | 783 ms | |
| Grok 4.6 | 6.4 s | 7 s | 10.9 s | |
| Grok 4.3 | 3.3 s | 3.7 s | 4 s | |
| Grok 4.20 (non-reasoning) | 439 ms | 457 ms | 502 ms | |
| Mistral Medium 3.5 | 306 ms | 344 ms | 428 ms | |
| Mistral Small 4 | 326 ms | 340 ms | 348 ms |
Asked again the same week
Week 36, 2026 was run 3 times, on September 2, 2026, September 2, 2026, September 2, 2026. Answers that changed between consecutive runs:
None. Every model gave the same majority answer each time.
Across editions
Week by week
One edition so far. Trends start with the second one; a line through a single point is a drawing, not a measurement.
Position changes
No model has changed its answer on a hamburger between consecutive editions.