En DashHotdogBenchmark

About

Every Monday we ask the biggest AI models whether a hot dog is a sandwich. Also a hamburger, also a taco. One word. Three times each, under three framings: asked plainly, told it is a sandwich, told it is not a sandwich. Then we publish exactly what they said, how long they took, and what it cost.

Why

The question is a joke. The plumbing is not. A real cross-provider benchmark has to deal with a different API per vendor, token counts that mean different things, latency that depends on where you run it, models that answer differently on Tuesday, and someone being down mid-run. This repo handles all of that in the smallest honest way and shows its work, so you can drop the hot dog and put your own question in its place.

The interesting number is not the answer. It is how much the answer moves when you tell the model what to think. That is a real property of a model, and it generalizes to every evaluation you would actually build.

How it is run

One run per week, filed by ISO week and kept forever. Re-running a week replaces that week's edition and keeps the old one under superseded/. No vendor is asked, sponsors this, or sees it first. The code, the raw data, and the methodology are all on GitHub.

Who

An En Dash Consulting project, planned and built with n-dx, our AI-powered development toolkit. The questions, the design, and the opinions are ours.

Corrections

Wrong data gets fixed at the source and republished, with the change in version control. If you think our classification rules are wrong, open an issue; that is a real argument and we want it.

Nothing here measures model quality. There is no right answer, and a model that says No is not better or worse than one that says Yes.