news
OpenArt Arena ranks video and image models by creative job, without publishing prompt, judge or vote counts
The September 15 launch splits rankings into boards for ads, film, animation, motion design, video editing and lip sync, judged blind by a named expert council plus a tastemaker pool still being recruited.
OpenArt announced a model leaderboard on September 15 that ranks image and video generators by professional use case instead of one overall score, and the launch materials omit the three numbers that would let an outsider judge how much evidence each ranking rests on: how many prompts, how many judges actually completed evaluations, and how many votes were cast.
Separate boards for lip sync, motion design and video editing
The video side of OpenArt Arena is split into a pooled overall ranking plus boards for ads, film, animation, motion design, video editing and lip sync, each with criteria specific to that kind of work. The image side splits into overall, e-commerce, film, graphic design and image editing. Judges see unlabeled outputs side by side and pick the one that best fulfills a given creative brief, scored on things like aesthetics, prompt adherence, realism and motion quality rather than a single technical figure. Those pairwise preferences are aggregated using the 1952 Bradley-Terry model, the same statistical machinery behind Elo-style leaderboards elsewhere. The launch page names Nano Banana Pro, GPT Image 2, Seedance 2.0 and Grok Imagine among the models ranked.
A named council, and a judging pool that is still being hired
The judging pool has two layers. The smaller Creative Expert Council includes named practitioners: Emmy-winning animation director William Lau, creative technologist Willonius Hatcher, marketing leader David Shing, and executives and educators from organizations including Edelman and UCLA. Around that, OpenArt says it is recruiting roughly 800 to 1,000 tastemakers from its own users, outside creative communities and working professionals in fields such as advertising. That range describes a planned pool, not people who have finished launch evaluations.
OpenArt has disclosed its statistical approach, its recruiting strategy, council names and an intention to publish some prompts. It has not disclosed prompt count, completed judge count or total votes, which is what would let anyone compare the scale of its launch evaluations against Contra Labs, which publishes evaluator count, task structure and total pairwise judgments, or Arena.ai, which shows vote counts and confidence ranges next to its rankings. OpenArt's announcement post also still carries an unresolved bracketed editorial note asking the company to confirm further methodology details and how often the leaderboard will update.
The first-of-its-kind framing has at least two prior examples
OpenArt describes Arena as the first leaderboard built on expert evaluations of top creatives across multiple industries. Contra Labs introduced its Human Creativity Benchmark in April 2026 using working creatives, real-world briefs, blinded model identities and Bradley-Terry aggregation into Elo-style rankings across landing pages, desktop apps, ad images, brand images and product videos. Artificial Analysis already ranks image models across a use-case taxonomy covering marketing, retail, live-action film, animation and others. What is genuinely new in OpenArt's version is the granularity of the professional boards, particularly lip sync and video editing, which the others do not break out.
The company running the leaderboard also sells the models on it
OpenArt's own product pages sell generation on Seedance 2.5, Kling 3.0, Sora 2, Seedance 2.0, WAN 2.7, LTX-2.3, HappyHorse, PixVerse and Gemini Omni Flash, so the Arena ranks the same models the platform monetizes. CEO and co-founder Coco Mao framed the project around the idea that creators "should not have to rely solely on technical benchmarks" when choosing tools. The company says it reaches over 8 million monthly active users, launched its Director video tool in June, and is part of the 2026 Disney Accelerator.
A reseller running a leaderboard is not automatically compromised, and use-case boards are a real improvement over one blended score for anyone deciding what to render tomorrow. But leaderboards in this category move fast and have been gamed by stealth entries before, and without vote and judge counts there is no way to tell whether a given board reflects a hundred considered comparisons or a handful.
Sources: OpenArt, Business Wire, OpenArt Arena, VentureBeat.
Daniel Ochoa: Covers model launches, shutdowns and pricing changes as they happen. Reads deprecation notices for a living so you do not have to.