Methodology
How we test
The battery
The prompts are not chosen to look good. They are chosen because they break models, and because a human can verify the result by eye: either the sign says OPEN 24H or it does not, either there are five fingers or there are not.
| Category | Prompts |
|---|---|
| Hands and fingers | 4 |
| Human faces in motion | 4 |
| Liquid physics | 4 |
| On-screen text | 4 |
| Multi-shot continuity | 4 |
| Lip-sync and dialogue | 4 |
| Camera movement | 4 |
| Animal anatomy | 4 |
| Style consistency | 4 |
| Image to video | 4 |
The battery is frozen. A new model gets exactly the battery every previous model got. If a prompt ever has to change, that creates a new battery version rather than silently rewriting history.
Scoring
A vision model evaluates each clip against the original prompt without being told which model produced it. It scores five dimensions from 0 to 5:
- Prompt fidelity — did it make what was asked for?
- Visual quality — resolution, artefacts, coherence.
- Physics and motion — does the world behave?
- Subject and scene consistency — does it hold together over time?
- Audio and sync — only where the model generates audio. Models without audio are not penalised; the dimension is simply excluded from their average.
Each clip is scored three separate times and the passes are averaged, which damps the judge's own variance. We review a 10% sample by hand to confirm the judge has not drifted.
Cost, and the number nobody publishes
Sticker price per generation is close to meaningless. What matters is how many attempts it took to get a clip you would actually use. We record that as attempts to usable and it is the figure that decides what a model really costs you.
What we do not do
- We do not accept payment for placement, ever.
- We do not blend in other people’s rankings and present the result as our own measurement.
- We do not publish a price we have not verified ourselves. Unverified third-party figures are labelled as such.
- We do not hide the failures. A model that fails a prompt gets a recorded failure, not a quiet retry until it works.
Conflicts of interest
Some links to platforms earn us a commission and are marked. Model makers do not pay us and cannot see results before publication. If the best-scoring model earns us nothing, it still finishes first. Full disclosure.