The Tuna Melt Inquiry, week by week
Open-faced and grilled. Where does the top slice go?
Week 36, 2026, up close
Did each model agree with itself?
Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.
| Model | Control | Asserted | Denied | Self-agreement |
|---|---|---|---|---|
| Claude Opus 5 | 100% | |||
| Claude Sonnet 5 | 100% | |||
| Claude Haiku 4.5 | 100% | |||
| GPT-5.6 Sol | 100% | |||
| GPT-5.5 | 100% | |||
| GPT-5.4 mini | 100% | |||
| Grok 4.6 | 100% | |||
| Grok 4.3 | 100% | |||
| Grok 4.20 (non-reasoning) | no data | 100% | ||
| Mistral Medium 3.5 | 100% | |||
| Mistral Small 4 | 100% | |||
| DeepSeek V4 Pro | 100% |
How much did latency swing?
The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.
| Model | Range | Fastest | Median | Slowest |
|---|---|---|---|---|
| Claude Opus 5 | 1.2 s | 1.4 s | 2.1 s | |
| Claude Sonnet 5 | 1.1 s | 1.2 s | 1.2 s | |
| Claude Haiku 4.5 | 620 ms | 645 ms | 767 ms | |
| GPT-5.6 Sol | 1 s | 1.1 s | 1.8 s | |
| GPT-5.5 | 961 ms | 1.1 s | 1.1 s | |
| GPT-5.4 mini | 741 ms | 801 ms | 1.5 s | |
| Grok 4.6 | 5.9 s | 6.8 s | 8.8 s | |
| Grok 4.3 | 2.7 s | 3.2 s | 3.7 s | |
| Grok 4.20 (non-reasoning) | 429 ms | 444 ms | 489 ms | |
| Mistral Medium 3.5 | 325 ms | 332 ms | 350 ms | |
| Mistral Small 4 | 301 ms | 311 ms | 376 ms | |
| DeepSeek V4 Pro | 1.6 s | 1.7 s | 2.1 s |
Asked again the same week
Week 36, 2026 was run 6 times, on September 2 and September 3, 2026. Answers that changed between consecutive runs:
None. Every model gave the same majority answer each time.
Across editions
Week by week
One edition so far. Trends start with the second one; a line through a single point is a drawing, not a measurement.
Position changes
No model has changed its answer on a tuna melt between consecutive editions.