SlopTV

Benchmark

Faces leaderboard

Every model runs the same 40 prompts. A vision model scores each clip blind against the original prompt in a single pass, across visual fidelity, physics and motion, subject consistency, prompt adherence and audio sync. Scores are out of 5, and the 95% CI column shows how much a score could plausibly move on a re-run with this few samples. Full methodology.

Board: Overall · Hands · Faces · Liquids · On-screen text · Multi-shot · Lip-sync · Camera · Animals · Style · Image to video

Clip shown: every model here on the same prompt, “face-turn-smile”, so the preview is a fair comparison, not a cherry-picked best clip.

#ModelClipOverall95% CIFidelity VisualPhysicsConsistency AudioAttemptsGen (s) $/runRuns
1 OpenAI Sora 2 discontinued
OpenAI
5.00 ±0 5.00 5.00 5.00 1.00 217 4
2 Google Veo 3.1
Google
5.00 ±0 5.00 5.00 5.00 132 4
3 MiniMax Hailuo 02
MiniMax
5.00 ±0 5.00 5.00 5.00 300 4
4 MiniMax H3
MiniMax
5.00 ±0 5.00 5.00 5.00 448 4
5 Seedance 2.0
ByteDance
5.00 ±0 5.00 5.00 5.00 212 3
6 Kling 3.0 Turbo
Kuaishou
4.97 ±0.1 4.95 4.95 4.95 5.00 5.00 4
7 Runway Gen-4.5
Runway
4.93 ±0.14 5.00 4.90 4.90 4.90 4
8 PixVerse v5.5
PixVerse
4.93 ±0.24 5.00 4.88 4.95 4.88 4
9 Kling 3.0
Kuaishou
4.79 ±0.5 4.75 4.75 4.88 1.00 190 4
10 Seedance 2.5
ByteDance
4.63 ±1.19 4.75 4.75 4.38 2.00 233 4
11 Grok Video
xAI
4.54 ±0.59 4.38 4.50 4.75 1.00 85 4

Attempts is the average number of generations needed before one was usable, the number that actually decides what a model costs you. 95% CI is computed from that model's own runs on this board, not assumed. On a category board with only 4 runs per model, a wide CI is the honest reflection of a small sample, not a bug. Two models whose CIs overlap are not meaningfully different at this sample size.

Frequently asked

Which AI video model is best at Faces right now?

As of 23 September 2026, OpenAI Sora 2 by OpenAI ranks #1 on Faces on SlopTV's benchmark, scoring 5.00/5 across 4 scored runs (±0 at 95% confidence).