Benchmark
Lip-sync leaderboard
Board: Overall · Hands · Faces · Liquids · On-screen text · Multi-shot · Lip-sync · Camera · Animals · Style · Image to video
Clip shown: every model here on the same prompt, “lipsync-direct-address”, so the preview is a fair comparison, not a cherry-picked best clip.
| # | Model | Clip | Overall | 95% CI | Fidelity | Visual | Physics | Consistency | Audio | Attempts | Gen (s) | $/run | Runs |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | OpenAI |
5.00 | ±0 | 5.00 | 5.00 | — | 5.00 | 5.00 | 1.00 | 217 | — | 4 | |
| 2 | xAI |
5.00 | ±0 | 5.00 | 5.00 | — | 5.00 | 5.00 | 1.00 | 85 | — | 4 | |
| 3 | ByteDance |
5.00 | ±0 | 5.00 | 5.00 | — | — | 5.00 | — | 263 | — | 4 | |
| 4 | ByteDance |
5.00 | ±0 | 5.00 | 5.00 | — | — | 5.00 | — | 233 | — | 4 | |
| 5 | MiniMax |
5.00 | ±0 | 5.00 | 5.00 | — | — | 5.00 | — | 448 | — | 4 | |
| 6 | Kuaishou |
4.87 | ±0.08 | 4.85 | 4.85 | 4.80 | 4.83 | 5.00 | — | — | — | 4 | |
| 7 | Runway |
4.78 | ±0 | 4.50 | 4.80 | 5.00 | 4.80 | — | — | — | — | 4 | |
| 8 | 4.50 | ±0.65 | 4.38 | 4.75 | — | — | 4.38 | — | 131 | — | 4 | ||
| 9 | PixVerse |
4.24 | ±0.71 | 3.50 | 4.53 | 4.63 | 4.28 | — | — | — | — | 4 | |
| 10 | Kuaishou |
3.77 | ±0.06 | 3.00 | 4.50 | — | 4.00 | — | 1.00 | 190 | — | 4 | |
| 11 | MiniMax |
0.83 | ±0.44 | 1.00 | 1.50 | — | — | 0.00 | — | 300 | — | 4 |
Attempts is the average number of generations needed before one was usable, the number that actually decides what a model costs you. 95% CI is computed from that model's own runs on this board, not assumed. On a category board with only 4 runs per model, a wide CI is the honest reflection of a small sample, not a bug. Two models whose CIs overlap are not meaningfully different at this sample size.
Frequently asked
Which AI video model is best at Lip-sync right now?
As of 23 September 2026, OpenAI Sora 2 by OpenAI ranks #1 on Lip-sync on SlopTV's benchmark, scoring 5.00/5 across 4 scored runs (±0 at 95% confidence).