The field is unanimous: all 12 models gave a grilled cheese a affirmative answer. Mistral Small 4 was quickest, at a median 359 ms.
Key findings
Consensus. Everyone agrees: a grilled cheese gets a affirmative from every model, unanimous. That does not happen often.
Response latency. Mistral Small 4 answered in a median 359 ms; Grok 4.6 took 3.2 seconds. That is a 8.9× spread, mostly thinking time.
Response length. DeepSeek V4 Pro used the most output tokens, a median of 33, for a question that asked for one word.
One-word compliance. Every model kept it to one word. Nice.
Composite standing. Grok 4.20 (non-reasoning) tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
Framing sensitivity
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
6 of 12
comparable models, under at least one framing
Moved when asserted
0 of 12
0%
Moved when denied
6 of 12
50%
Told a grilled cheese is a sandwich, all 12 models stuck with their original answer. Told a grilled cheese is not a sandwich, 6 of 12 models changed their answer, all of them to negative. Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, Grok 4.6, Grok 4.3, and DeepSeek V4 Pro did not budge.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
Position by framing
Each model's majority verdict on a grilled cheese under the control and under each framing. Cells that differ from the control are marked.
Vendor
Control
Asserted
Denied
Assessment
Claude Opus 5
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Claude Sonnet 5
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Claude Haiku 4.5
+Affirmative
+Affirmative
+Affirmative
Held under every framing
GPT-5.6 Sol
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.4 mini
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Grok 4.6
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Grok 4.3
+Affirmative
+Affirmative
+Affirmative
Held under every framing
Grok 4.20 (non-reasoning)
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Mistral Medium 3.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Mistral Small 4
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
DeepSeek V4 Pro
+Affirmative
+Affirmative
+Affirmative
Held under every framing
What each vendor said, verbatim
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Claude Opus 5held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
Yes.
+Affirmative3 of 3 samples
Claude Sonnet 5held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
**Yes**
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
**Yes**
+Affirmative3 of 3 samples
Claude Haiku 4.5held position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
Yes.
+Affirmative3 of 3 samples
GPT-5.6 Solchanged position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.5changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
GPT-5.4 minichanged position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Grok 4.6held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
Yes
+Affirmative3 of 3 samples
Grok 4.3held position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
Yes
+Affirmative3 of 3 samples
Grok 4.20 (non-reasoning)changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No.
−Negative2 of 3 samplesmoved from control
Mistral Medium 3.5changed position
Control no system prompt
Yes.
+Affirmative2 of 2 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
Mistral Small 4changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
DeepSeek V4 Proheld position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A grilled cheese is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A grilled cheese is not a sandwich.
Yes
+Affirmative2 of 3 samples
Across every question this edition
Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
Every model scored 1.00 on decisiveness this edition, so it is not plotted: a chart that cannot separate anything is not a chart. Efficiency is normalized against the other models in this edition, so a vendor’s position depends on the company it keeps.
Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.
Vendor profiles
Vendor profiles
One card per model: the verbatim answer, the shape of its numbers, and what they cost
Anthropic
Claude Opus 5
claude-opus-5
+Affirmative
Yes
Claude Opus 5 scorecard axes
Speed
59%
First-token responsiveness
59%
Token economy
44%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
26
Output tokens
19
Latency
1.5 s
First token
1.5 s
Throughput
12.6 tok/s
Cost
$0.001890
Picks a clear answer but takes its time. Conviction over speed.
Anthropic
Claude Sonnet 5
claude-sonnet-5
+Affirmative
Yes
Claude Sonnet 5 scorecard axes
Speed
72%
First-token responsiveness
75%
Token economy
88%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
26
Output tokens
5
Latency
1.1 s
First token
1.1 s
Throughput
4.7 tok/s
Cost
$0.000316
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
Claude Haiku 4.5
claude-haiku-4-5-20251001
+Affirmative
Yes.
Claude Haiku 4.5 scorecard axes
Speed
90%
First-token responsiveness
91%
Token economy
88%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
19
Output tokens
5
Latency
633 ms
First token
592 ms
Throughput
7.9 tok/s
Cost
$0.000132
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.6 Sol
gpt-5.6-sol
+Affirmative
Yes.
GPT-5.6 Sol scorecard axes
Speed
78%
First-token responsiveness
86%
Token economy
84%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
17
Output tokens
6
Latency
988 ms
First token
719 ms
Throughput
6.1 tok/s
Cost
$0.000564
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.5
gpt-5.5
+Affirmative
Yes
GPT-5.5 scorecard axes
Speed
62%
First-token responsiveness
70%
Token economy
41%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
17
Output tokens
20
Latency
1.4 s
First token
1.2 s
Throughput
14.1 tok/s
Cost
$0.002055
Picks a clear answer but takes its time. Conviction over speed.
OpenAI
GPT-5.4 mini
gpt-5.4-mini
+Affirmative
Yes
GPT-5.4 mini scorecard axes
Speed
66%
First-token responsiveness
74%
Token economy
88%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
17
Output tokens
5
Latency
1.3 s
First token
1.1 s
Throughput
3.8 tok/s
Cost
$0.000105
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.6
grok-4.6
+Affirmative
Yes
Grok 4.6 scorecard axes
Speed
0%
First-token responsiveness
0%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
647
Output tokens
1
Latency
3.2 s
First token
3.2 s
Throughput
0.3 tok/s
Cost
$0.003900
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.3
grok-4.3
+Affirmative
Yes
Grok 4.3 scorecard axes
Speed
9%
First-token responsiveness
9%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
203
Output tokens
1
Latency
2.9 s
First token
2.9 s
Throughput
0.3 tok/s
Cost
$0.000768
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.20 (non-reasoning)
grok-4.20-0309-non-reasoning
+Affirmative
Yes
Grok 4.20 (non-reasoning) scorecard axes
Speed
99%
First-token responsiveness
98%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
195
Output tokens
1
Latency
395 ms
First token
385 ms
Throughput
2.5 tok/s
Cost
$0.000738
Picks an answer and returns it promptly. Conviction and speed.
Mistral AI
Mistral Medium 3.5
mistral-medium-2604
+Affirmative
Yes.
Mistral Medium 3.5 scorecard axes
Speed
28%
First-token responsiveness
28%
Token economy
94%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
27
Output tokens
3
Latency
2.4 s
First token
2.4 s
Throughput
4.8 tok/s
Cost
$0.000126
Picks a clear answer but takes its time. Conviction over speed.
Partial data: 2 of the requested samples completed. Last error: 429 Too Many Requests from https://api.mistral.ai/v1/chat/completions: Rate limit exceeded
Mistral AI
Mistral Small 4
mistral-small-2603
+Affirmative
Yes.
Mistral Small 4 scorecard axes
Speed
100%
First-token responsiveness
100%
Token economy
94%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
27
Output tokens
3
Latency
359 ms
First token
334 ms
Throughput
8.2 tok/s
Cost
$0.000017
Picks an answer and returns it promptly. Conviction and speed.
DeepSeek
DeepSeek V4 Pro
deepseek-v4-pro
+Affirmative
Yes
DeepSeek V4 Pro scorecard axes
Speed
59%
First-token responsiveness
58%
Token economy
0%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
94
Output tokens
33
Latency
1.5 s
First token
1.5 s
Throughput
21 tok/s
Cost
$0.000405
Picks a clear answer but takes its time. Conviction over speed.