SlopTV

news

A public video model leaderboard is ranking 12 models on 1,394 blind votes total

The LLM Stats video arena, refreshed September 20, puts Kling v3 first at 1,934 ahead of Happy Horse 1.0 and Seedance 2.0 Fast, on a vote pool reported as a single number spanning three different tasks.

By Priya Shenoy ยท

A public video model leaderboard run by LLM Stats, refreshed on September 20, 2026, ranks 12 generation models using 1,394 blind votes in total, spread across text-to-video, image-to-video and video editing. Kling v3 leads with a score of 1,934, ahead of Happy Horse 1.0 at 1,816 and Seedance 2.0 Fast at 1,747.

A public AI video leaderboard refreshed on September 20, 2026 ranks twelve generation models off a total of 1,394 blind votes, and that total covers three different tasks at once. The ordering it produces is not implausible, but the vote pool is small enough that the distance between any two adjacent models deserves more caution than a ranked list tends to invite.

Kling v3 at 1,934, Happy Horse 1.0 at 1,816, Seedance 2.0 Fast at 1,747

The LLM Stats video arena reports Kling v3 in first place on text-to-video with an arena score of 1,934. Happy Horse 1.0 sits second at 1,816, and Seedance 2.0 Fast third at 1,747. The gap between first and second is 118 points, and between second and third it is 69 points. Each figure is presented as a single number in the summary, with no interval attached to it.

One aggregate vote count covering three separate tasks

The leaderboard states that its rankings come from 1,394 blind votes across text-to-video, image-to-video and video editing, with voters comparing real outputs without being told which model produced which clip. Twelve models are listed as reviewed. Those two numbers are the ones to hold onto, because the page reports the vote count as one aggregate rather than breaking it out by task.

The arithmetic is worth doing. If every vote is a pairwise comparison, 1,394 votes produce roughly 2,788 model appearances, which averages to about 232 appearances per model across the whole board. Split evenly across the three task categories, that falls to something in the range of 75 to 80 comparisons per model per task. Votes are almost certainly not distributed evenly, since popular models attract more matchups than obscure ones, so the busiest entries will sit above that average and the quietest ones well below it. Nothing on the summary indicates which models are which.

Why the gaps between adjacent models are hard to read

Elo-style ratings converge slowly. At a few dozen to a couple of hundred comparisons per model, the uncertainty around a rating is typically large enough to swallow differences of this size, which means a 69 point gap between second and third place may reflect the sampling as much as it reflects the models. The leaderboard does not surface confidence bounds next to the top-line scores, so a reader has no way to tell from the ranking alone whether Happy Horse 1.0 and Seedance 2.0 Fast are genuinely separated or simply ordered.

What the page does disclose, and what it does not

To its credit, the arena publishes a vote count and a methodology link at all, which is more than several creative-model rankings bother with, including the OpenArt Arena creative job leaderboard. The ingredients that are still missing from the summary view are the per-task vote split, the per-model matchup counts, the prompt set, and any statement on how repeat voters or automated traffic are handled. Those four things are what separate a ranking you can act on from a ranking you can only read.

For anyone choosing a model on the strength of a leaderboard screenshot, the practical takeaway is narrow: Kling v3 is ahead on this particular board today, by a margin drawn from a vote pool of under 1,400.

Sources: LLM Stats.


PS

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.