SlopTV

Benchmark

Image to video leaderboard

Every model runs the same 40 prompts. A vision model scores each clip blind against the original prompt in a single pass, across visual fidelity, physics and motion, subject consistency, prompt adherence and audio sync. Scores are out of 5, and the 95% CI column shows how much a score could plausibly move on a re-run with this few samples. Full methodology.

Board: Overall · Hands · Faces · Liquids · On-screen text · Multi-shot · Lip-sync · Camera · Animals · Style · Image to video

Clip shown: every model here on the same prompt, “i2v-portrait-animate”, so the preview is a fair comparison, not a cherry-picked best clip.

#ModelClipOverall95% CIFidelity VisualPhysicsConsistency AudioAttemptsGen (s) $/runRuns
1 OpenAI Sora 2 discontinued
OpenAI
5.00 ±0 5.00 5.00 1.00 217 4
2 Kling 3.0
Kuaishou
5.00 ±0 5.00 5.00 1.00 190 4
3 Grok Video
xAI
5.00 ±0 5.00 5.00 1.00 85 4
4 Seedance 2.0
ByteDance
5.00 ±0 5.00 5.00 292 4
5 Seedance 2.5
ByteDance
5.00 ±0 5.00 5.00 233 4
6 Google Veo 3.1
Google
5.00 ±0 5.00 5.00 75 4
7 MiniMax Hailuo 02
MiniMax
5.00 ±0 5.00 5.00 300 4
8 MiniMax H3
MiniMax
5.00 ±0 5.00 5.00 448 4
9 Kling 3.0 Turbo
Kuaishou
5.00 ±0 5.00 5.00 5.00 5.00 5.00 4
10 PixVerse v5.5
PixVerse
4.93 ±0.21 4.88 4.95 4.95 4.95 4
11 Runway Gen-4.5
Runway
4.59 ±0.76 4.70 4.50 4.58 4.58 4

Attempts is the average number of generations needed before one was usable, the number that actually decides what a model costs you. 95% CI is computed from that model's own runs on this board, not assumed. On a category board with only 4 runs per model, a wide CI is the honest reflection of a small sample, not a bug. Two models whose CIs overlap are not meaningfully different at this sample size.

Frequently asked

Which AI video model is best at Image to video right now?

As of 23 September 2026, OpenAI Sora 2 by OpenAI ranks #1 on Image to video on SlopTV's benchmark, scoring 5.00/5 across 4 scored runs (±0 at 95% confidence).