En DashHotdogBenchmark

Week 36, 2026 edition

An archived record of every question asked in this edition. Current reports are on the home page.

The Hot Dog Question

Full archived report →

Key performance indicators

  • Models evaluated

    11

  • Consensus position

    Affirmative

    55% of the field

  • Median latency

    1.3 s

  • Median output tokens

    5

  • Instruction compliance

    97%

Executive summary

A affirmative answer on a hot dog has majority support this week: 6 of 11 models (55%). Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 disagree. Mistral Small 4 was quickest, at a median 347 ms. 97% of answers were actually one word, as asked.

Key findings

  • Consensus. 55% of models (6 of 11) say affirmative.
  • Dissent. Claude Opus 5, Grok 4.6, Grok 4.3, Mistral Medium 3.5, and Mistral Small 4 went the other way.
  • Response latency. Mistral Small 4 answered in a median 347 ms; Grok 4.6 took 12.6 seconds. That is a 36.3× spread, mostly thinking time.
  • Response length. Claude Opus 5 used the most output tokens, a median of 93, for a question that asked for one word.
  • Instruction compliance. Claude Sonnet 5 did not always keep it to one word. Following the instruction and picking an answer are scored separately; they are different skills.
  • Composite standing. Mistral Small 4 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.

The Hamburger Control

Full archived report →

Key performance indicators

  • Models evaluated

    11

  • Consensus position

    Affirmative

    100% of the field

  • Median latency

    1.1 s

  • Median output tokens

    5

  • Instruction compliance

    100%

Executive summary

The field is unanimous: all 11 models gave a hamburger a affirmative answer. Mistral Small 4 was quickest, at a median 340 ms.

Key findings

  • Consensus. Everyone agrees: a hamburger gets a affirmative from every model, unanimous. That does not happen often.
  • Response latency. Mistral Small 4 answered in a median 340 ms; Grok 4.6 took 7.0 seconds. That is a 20.8× spread, mostly thinking time.
  • Response length. Claude Opus 5 used the most output tokens, a median of 123, for a question that asked for one word.
  • Instruction compliance. Every model kept it to one word. Nice.
  • Composite standing. Mistral Small 4 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.

The Taco Boundary

Full archived report →

Key performance indicators

  • Models evaluated

    11

  • Consensus position

    Negative

    100% of the field

  • Median latency

    940 ms

  • Median output tokens

    3

  • Instruction compliance

    100%

Executive summary

The field is unanimous: all 11 models gave a taco a negative answer. Mistral Medium 3.5 was quickest, at a median 365 ms.

Key findings

  • Consensus. Everyone agrees: a taco gets a negative from every model, unanimous. That does not happen often.
  • Response latency. Mistral Medium 3.5 answered in a median 365 ms; Grok 4.6 took 10.5 seconds. That is a 28.7× spread, mostly thinking time.
  • Response length. Claude Opus 5 used the most output tokens, a median of 68, for a question that asked for one word.
  • Instruction compliance. Every model kept it to one word. Nice.
  • Composite standing. Mistral Medium 3.5 tops the composite score at 1.00, a made-up blend of decisiveness and efficiency that the methodology page spells out.