The Taco Boundary, week by week
The only sandwich question with case law. Massachusetts, 2006.
Week 36, 2026, up close
Did each model agree with itself?
Every sample's answer, in the order it was taken, under each framing. A row of identical chips is a model that has made up its mind. Y is yes, N is no, ~ is a hedge.
| Model | Control | Asserted | Denied | Self-agreement |
|---|---|---|---|---|
| Claude Opus 5 | 89% | |||
| Claude Sonnet 5 | 100% | |||
| Claude Haiku 4.5 | 100% | |||
| GPT-5.6 Sol | 100% | |||
| GPT-5.5 | 100% | |||
| GPT-5.4 mini | 100% | |||
| Grok 4.6 | 100% | |||
| Grok 4.3 | 100% | |||
| Grok 4.20 (non-reasoning) | 89% | |||
| Mistral Medium 3.5 | 100% | |||
| Mistral Small 4 | 100% |
How much did latency swing?
The fastest and slowest call behind each median, asked plainly. A long bar is a model whose speed you cannot count on.
| Model | Range | Fastest | Median | Slowest |
|---|---|---|---|---|
| Claude Opus 5 | 2 s | 2.1 s | 3.4 s | |
| Claude Sonnet 5 | 877 ms | 940 ms | 1 s | |
| Claude Haiku 4.5 | 548 ms | 617 ms | 705 ms | |
| GPT-5.6 Sol | 1 s | 1.2 s | 1.5 s | |
| GPT-5.5 | 1.1 s | 1.1 s | 1.9 s | |
| GPT-5.4 mini | 496 ms | 519 ms | 580 ms | |
| Grok 4.6 | 6.2 s | 10.5 s | 11.2 s | |
| Grok 4.3 | 3.9 s | 4.1 s | 4.5 s | |
| Grok 4.20 (non-reasoning) | 369 ms | 448 ms | 477 ms | |
| Mistral Medium 3.5 | 321 ms | 365 ms | 371 ms | |
| Mistral Small 4 | 341 ms | 414 ms | 637 ms |
Asked again the same week
Week 36, 2026 was run 3 times, on September 2, 2026, September 2, 2026, September 2, 2026. Answers that changed between consecutive runs:
None. Every model gave the same majority answer each time.
Across editions
Week by week
One edition so far. Trends start with the second one; a line through a single point is a drawing, not a measurement.
Position changes
No model has changed its answer on a taco between consecutive editions.