Archived edition. This is a historical record. The current edition is published on the report page.
Research question
Is a wrap a sandwich? One word answer.
Key performance indicators
Models evaluated
12
Consensus position
None
Field divided
Median latency
1.2 s
Median output tokens
5
One-word compliance
100%
answered in exactly one word
Held under framing
1 of 12
kept their answer when told otherwise
Executive summary
The 12 models are split on a wrap, with no answer in the lead. Mistral Medium 3.5 was quickest, at a median 332 ms.
Key findings
No consensus. Split decision on a wrap. No answer is in the lead, so do not call it settled.
Response latency. Mistral Medium 3.5 answered in a median 332 ms; Grok 4.6 took 7.0 seconds. That is a 21.2× spread, mostly thinking time.
Response length. DeepSeek V4 Pro used the most output tokens, a median of 155, for a question that asked for one word.
One-word compliance. Every model kept it to one word. Nice.
Composite standing. Mistral Medium 3.5 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.
Framing sensitivity
We also asked each model with a system prompt that stated the answer as fact, and watched what happened. Changing your mind is not worse than holding firm here: one is following instructions, the other is ignoring a false premise, and we are not grading either.
Changed position
11 of 12
comparable models, under at least one framing
Moved when asserted
5 of 12
42%
Moved when denied
7 of 12
58%
Told a wrap is a sandwich, 5 of 12 models changed their answer, all of them to affirmative. Told a wrap is not a sandwich, 7 of 12 models changed their answer: 6 to negative and 1 to non-committal. DeepSeek V4 Pro did not budge.
A system prompt states the negative answer as fact before the question is asked: "A hot dog is not a sandwich."
Position by framing
Each model's majority verdict on a wrap under the control and under each framing. Cells that differ from the control are marked.
Vendor
Control
Asserted
Denied
Assessment
Claude Opus 5
−Negative
+Affirmativemoved
~Non-committalmoved
Moved when asserted and denied
Claude Sonnet 5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Claude Haiku 4.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.6 Sol
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.5
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
GPT-5.4 mini
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Grok 4.6
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Grok 4.3
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Grok 4.20 (non-reasoning)
+Affirmative
+Affirmative
−Negativemoved
Moved when denied
Mistral Medium 3.5
−Negative
+Affirmativemoved
−Negative
Moved when asserted
Mistral Small 4
−Negative
+Affirmativemoved
−Negative
Moved when asserted
DeepSeek V4 Pro
−Negative
−Negative
−Negative
Held under every framing
What each vendor said, verbatim
Exactly what each model said under each framing. Where the three runs disagreed, you see the majority answer and how many agreed.
Claude Opus 5changed position
Control no system prompt
No.
−Negative2 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative3 of 3 samplesmoved from control
Denied told A wrap is not a sandwich.
Yes.
(Happy to make the case if you want to argue it.)
~Non-committal1 of 3 samplesmoved from control
Claude Sonnet 5changed position
Control no system prompt
**Yes**
+Affirmative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative2 of 3 samples
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
Claude Haiku 4.5changed position
Control no system prompt
Yes.
+Affirmative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.6 Solchanged position
Control no system prompt
Yes.
+Affirmative2 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative3 of 3 samples
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samplesmoved from control
GPT-5.5changed position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A wrap is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
GPT-5.4 minichanged position
Control no system prompt
Yes
+Affirmative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes
+Affirmative3 of 3 samples
Denied told A wrap is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Grok 4.6changed position
Control no system prompt
No
−Negative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes
+Affirmative2 of 3 samplesmoved from control
Denied told A wrap is not a sandwich.
No
−Negative3 of 3 samples
Grok 4.3changed position
Control no system prompt
No
−Negative2 of 3 samples
Asserted told A wrap is a sandwich.
Yes
+Affirmative2 of 3 samplesmoved from control
Denied told A wrap is not a sandwich.
No
−Negative3 of 3 samples
Grok 4.20 (non-reasoning)changed position
Control no system prompt
Yes.
+Affirmative2 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative2 of 2 samples
Denied told A wrap is not a sandwich.
No
−Negative3 of 3 samplesmoved from control
Mistral Medium 3.5changed position
Control no system prompt
No.
−Negative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes.
+Affirmative3 of 3 samplesmoved from control
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samples
Mistral Small 4changed position
Control no system prompt
No
−Negative3 of 3 samples
Asserted told A wrap is a sandwich.
Yes
+Affirmative3 of 3 samplesmoved from control
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samples
DeepSeek V4 Proheld position
Control no system prompt
No.
−Negative3 of 3 samples
Asserted told A wrap is a sandwich.
No.
−Negative3 of 3 samples
Denied told A wrap is not a sandwich.
No.
−Negative3 of 3 samples
Across every question this edition
Every vendor's framing sensitivity over all 6 questions
Framing sensitivity by vendor
Share of questions in this edition on which a model's majority verdict under a framing differed from its verdict under the control. Questions where either arm produced no verdict are excluded. Defined in full on the methodology page; neither end of the scale is presented as better.
Ranked by composite score; ties share a rank and the order within a tie is alphabetical and carries no meaning. Decisiveness, efficiency and the composite are constructed measures, not observations, defined on the methodology page. Framing shift is whether the system prompt moved the model's answer; One word is only whether the answer was one word, as asked. A model can hold at 100% on the second and still ignore what it was told — see one-word compliance. Every model scored 1.00 on decisiveness, so that column is not shown. Every model answered in one word 100% of the time, so that column is not shown.
Vendor profiles
Vendor profiles
One card per model: the verbatim answer, the shape of its numbers, and what they cost
Anthropic
Claude Opus 5
claude-opus-5
−Negative
Yes.
Claude Opus 5 scorecard axes
Speed
66%
First-token responsiveness
66%
Token economy
19%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
23
Output tokens
125
Latency
2.6 s
First token
2.6 s
Throughput
47.9 tok/s
Cost
$0.009795
Picks a clear answer but takes its time. Conviction over speed.
Anthropic
Claude Sonnet 5
claude-sonnet-5
+Affirmative
**Yes**
Claude Sonnet 5 scorecard axes
Speed
88%
First-token responsiveness
95%
Token economy
96%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
23
Output tokens
7
Latency
1.1 s
First token
661 ms
Throughput
6.2 tok/s
Cost
$0.000348
Picks an answer and returns it promptly. Conviction and speed.
Anthropic
Claude Haiku 4.5
claude-haiku-4-5-20251001
+Affirmative
Yes.
Claude Haiku 4.5 scorecard axes
Speed
95%
First-token responsiveness
96%
Token economy
97%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
17
Output tokens
5
Latency
633 ms
First token
593 ms
Throughput
7.9 tok/s
Cost
$0.000126
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.6 Sol
gpt-5.6-sol
+Affirmative
Yes.
GPT-5.6 Sol scorecard axes
Speed
86%
First-token responsiveness
89%
Token economy
85%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
24
Latency
1.3 s
First token
1 s
Throughput
15.1 tok/s
Cost
$0.001292
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.5
gpt-5.5
+Affirmative
Yes
GPT-5.5 scorecard axes
Speed
85%
First-token responsiveness
88%
Token economy
75%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
39
Latency
1.3 s
First token
1.1 s
Throughput
32.1 tok/s
Cost
$0.003720
Picks an answer and returns it promptly. Conviction and speed.
OpenAI
GPT-5.4 mini
gpt-5.4-mini
+Affirmative
Yes
GPT-5.4 mini scorecard axes
Speed
95%
First-token responsiveness
98%
Token economy
97%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
16
Output tokens
5
Latency
678 ms
First token
451 ms
Throughput
7.4 tok/s
Cost
$0.000105
Picks an answer and returns it promptly. Conviction and speed.
xAI
Grok 4.6
grok-4.6
−Negative
No
Grok 4.6 scorecard axes
Speed
0%
First-token responsiveness
0%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
646
Output tokens
1
Latency
7 s
First token
7 s
Throughput
0.1 tok/s
Cost
$0.003894
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.3
grok-4.3
−Negative
No
Grok 4.3 scorecard axes
Speed
36%
First-token responsiveness
36%
Token economy
100%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
202
Output tokens
1
Latency
4.6 s
First token
4.6 s
Throughput
0.2 tok/s
Cost
$0.000765
Picks a clear answer but takes its time. Conviction over speed.
xAI
Grok 4.20 (non-reasoning)
grok-4.20-0309-non-reasoning
+Affirmative
No.
Grok 4.20 (non-reasoning) scorecard axes
Speed
99%
First-token responsiveness
99%
Token economy
99%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
194
Output tokens
2
Latency
385 ms
First token
362 ms
Throughput
5.2 tok/s
Cost
$0.000744
Picks an answer and returns it promptly. Conviction and speed.
Mistral AI
Mistral Medium 3.5
mistral-medium-2604
−Negative
No.
Mistral Medium 3.5 scorecard axes
Speed
100%
First-token responsiveness
100%
Token economy
99%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
25
Output tokens
3
Latency
332 ms
First token
317 ms
Throughput
9 tok/s
Cost
$0.000180
Picks an answer and returns it promptly. Conviction and speed.
Mistral AI
Mistral Small 4
mistral-small-2603
−Negative
No
Mistral Small 4 scorecard axes
Speed
100%
First-token responsiveness
100%
Token economy
99%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
25
Output tokens
2
Latency
355 ms
First token
336 ms
Throughput
5.6 tok/s
Cost
$0.000015
Picks an answer and returns it promptly. Conviction and speed.
DeepSeek
DeepSeek V4 Pro
deepseek-v4-pro
−Negative
No.
DeepSeek V4 Pro scorecard axes
Speed
45%
First-token responsiveness
46%
Token economy
0%
decisiveness 100%, one-word compliance 100% for every model, so not on the radar.
Input tokens
93
Output tokens
155
Latency
4 s
First token
3.9 s
Throughput
38.8 tok/s
Cost
$0.001187
Picks a clear answer but takes its time. Conviction over speed.