study
Where AI video still breaks: motion over time, not single frames
Faces, style and image-to-video are close to solved in our tests. Melting ice, a full camera orbit and a galloping horse are not.
A single beautiful frame is no longer impressive. Every model at the top of our leaderboard can produce one. What separates them, and what decides whether generated video can be used in real production, is whether a scene stays coherent while it moves.
What is close to solved
Averaged across the 12 models we have scored, some categories are near the ceiling:
| Category | Average score |
|---|---|
| Image to video | 4.96 |
| Human faces | 4.89 |
| Multi-shot continuity | 4.87 |
| Style | 4.76 |
| On-screen text | 4.74 |
Faces and image-to-video were the headline failures of earlier generations. On our prompts, they are now the safe ground.
What still breaks
The lowest-scoring prompts in the battery have one thing in common: something has to change over time while staying physically consistent.
| Prompt | Average | Models below 4 |
|---|---|---|
| Ice melting in whisky | 3.47 | 7 of 12 |
| Full 360-degree orbit | 3.97 | 5 of 12 |
| Galloping horse | 4.01 | 6 of 12 |
| Chopping a pepper | 4.18 | 5 of 12 |
| Cat jumping to a counter | 4.20 | 5 of 12 |
The failure modes repeat across models. Ice renders as a still image and never melts; in one clip the cube is identical between frames two seconds apart. A stream of liquid appears out of thin air and pools outside the glass. A camera asked to circle a seated reader stops partway round, and a bookshelf moves to a different wall mid-shot. Pepper slices duplicate or appear under the knife before it cuts, and in one clip a third hand joins in. A horse's hooves float and slide, or blur at exactly the moment they should meet the ground.
These are failures of temporal coherence: the model keeps each frame plausible but loses track of the scene between frames. It is also where specialist research effort is concentrated, because it is what stops generated video from replacing a shoot for anything that moves.
Who handles it best
The orbit splits the field cleanly. Six models, including Kling 3.0, Seedance 2.0 and Gemini Omni, complete the circle with the room holding together. The rest fall well short: Seedance 2.5 manages a 30 to 45 degree pan, Sora 2 about 60 to 90 degrees, Veo 3.1 about 270 degrees with the room rearranging itself, and both MiniMax models barely move the camera. So a model that ranks high overall can still fail the one shot your brief depends on. On hands, category scores range from 3.46 to 5.00 depending on the model, so the hands board is worth checking before you pick a model for anything with close-up handwork.
What the benchmark does not cover
Production teams talk about consistency as the binding constraint, and they mean something bigger than one clip: the same presenter, product or brand look held identical across dozens of separate shots, angles and languages. Our battery tests continuity inside a clip and across the cuts of a multi-shot prompt, not across a campaign of separate generations. A high multi-shot score is a good sign, not proof that a model will keep a character on-model across 50 clips. How we score is in the methodology.
Frequently asked
What is the hardest thing for AI video models in 2026?
Keeping physics and objects coherent while things move. In our benchmark the lowest-scoring prompts are ice melting in a drink, a full camera orbit, a galloping horse, chopping vegetables and a cat jumping onto a counter.
Are AI video models good at consistent characters now?
Within a single clip, mostly yes: our multi-shot category averages 4.87 out of 5 and faces 4.89. Keeping the same character identical across dozens of separate clips and languages is a harder, production-level problem our battery does not fully test.
Mara Voss: Runs the prompt battery and scores every clip before it is published. Came to AI video testing from QA automation, where she spent years building test suites designed to break things on purpose.