SlopTV

Benchmark

Lip-sync leaderboard

Every model runs the same 40 prompts. A vision model scores each clip blind against the original prompt in a single pass, across visual fidelity, physics and motion, subject consistency, prompt adherence and audio sync. Scores are out of 5, and the 95% CI column shows how much a score could plausibly move on a re-run with this few samples. Full methodology.

Board: Overall · Hands · Faces · Liquids · On-screen text · Multi-shot · Lip-sync · Camera · Animals · Style · Image to video

Clip shown: every model here on the same prompt, “lipsync-direct-address”, so the preview is a fair comparison, not a cherry-picked best clip.

#ModelClipOverall95% CIFidelity VisualPhysicsConsistency AudioAttemptsGen (s) $/runRuns
1 OpenAI Sora 2 discontinued
OpenAI
5.00 ±0 5.00 5.00 5.00 5.00 1.00 217 4
2 Grok Video
xAI
5.00 ±0 5.00 5.00 5.00 5.00 1.00 85 4
3 Seedance 2.0
ByteDance
5.00 ±0 5.00 5.00 5.00 263 4
4 Seedance 2.5
ByteDance
5.00 ±0 5.00 5.00 5.00 233 4
5 MiniMax H3
MiniMax
5.00 ±0 5.00 5.00 5.00 448 4
6 Kling 3.0 Turbo
Kuaishou
4.87 ±0.08 4.85 4.85 4.80 4.83 5.00 4
7 Runway Gen-4.5
Runway
4.78 ±0 4.50 4.80 5.00 4.80 4
8 Google Veo 3.1
Google
4.50 ±0.65 4.38 4.75 4.38 131 4
9 PixVerse v5.5
PixVerse
4.24 ±0.71 3.50 4.53 4.63 4.28 4
10 Kling 3.0
Kuaishou
3.77 ±0.06 3.00 4.50 4.00 1.00 190 4
11 MiniMax Hailuo 02
MiniMax
0.83 ±0.44 1.00 1.50 0.00 300 4

Attempts is the average number of generations needed before one was usable, the number that actually decides what a model costs you. 95% CI is computed from that model's own runs on this board, not assumed. On a category board with only 4 runs per model, a wide CI is the honest reflection of a small sample, not a bug. Two models whose CIs overlap are not meaningfully different at this sample size.

Frequently asked

Which AI video model is best at Lip-sync right now?

As of 23 September 2026, OpenAI Sora 2 by OpenAI ranks #1 on Lip-sync on SlopTV's benchmark, scoring 5.00/5 across 4 scored runs (±0 at 95% confidence).