SlopTV

news

An image-generation leaderboard refreshed September 28 ranks 10 models on 11,311 blind votes total

LLM Stats puts GPT Image 2 first with an arena score of 506, ahead of Grok Imagine Image 2.0 at 362 and MAI-Image-2.5 at 257, on a TrueSkill rating that penalises under-sampled models without publishing per-model vote counts.

By Priya Shenoy ยท

An LLM Stats image-generation leaderboard updated September 28, 2026 ranks 10 models using 11,311 blind human votes across text-to-image and image editing. GPT Image 2 leads with an arena score of 506, followed by Grok Imagine Image 2.0 at 362 and MAI-Image-2.5 at 257. Per-model vote counts are not published.

An image-generation leaderboard updated on September 28, 2026 ranks its entire field on 11,311 blind human votes, spread across 10 models and two different task types. That works out to somewhere near a thousand comparisons per model before the split between text-to-image and image editing is applied, and the site does not publish how those votes divide, either by model or by arena.

GPT Image 2 leads at 506, with a 144-point gap to second place

The published ordering has GPT Image 2 from OpenAI in first place with an arena score of 506, Grok Imagine Image 2.0 second at 362, and MAI-Image-2.5 third at 257. The rankings are described as coming from blind human votes across text-to-image generation and image editing tasks, where voters compare real outputs without knowing which model produced them.

The gap between first and second is 144 points. Whether that gap means GPT Image 2 makes better pictures, or simply that it has been shown to voters more often, depends entirely on the scoring method.

The scoring subtracts three standard deviations from every model

The board states its method plainly, which is more than several competing leaderboards do. Rankings use a TrueSkill conservative rating, defined as the mean skill estimate minus three standard deviations of uncertainty. Each prompt generates four images from randomly sampled models, voters see them side by side, and they pick the best and the worst without seeing model names or providers, a design the site says eliminates brand bias.

That conservative rating is the part worth understanding before reading the numbers as quality scores. Subtracting three standard deviations means a model with few recorded comparisons carries a wide uncertainty band and gets pushed down the table for that reason alone, independently of how good its output looked to the people who did see it. At roughly a thousand votes per model across two task types, uncertainty is doing real work in the ordering, and newer or less-sampled entrants are structurally disadvantaged. The formula is a defensible choice, and it is also the reason the ranking cannot be read as a straight quality ladder.

What is missing is the per-model sample size

The single number that would let a reader separate quality from sampling, the vote count behind each individual model, is not published. Neither is the prompt set, the split of the 11,311 votes between the generation arena and the editing arena, nor any information about who the voters are or whether repeat voting from the same person is controlled. With a conservative rating formula that is explicitly sensitive to sample size, withholding sample size per model leaves the most important variable unreadable.

Eleven thousand votes is small, and also not the smallest

For scale, Arena's image-to-video board reported 2,081,953 votes across 48 models as of September 21, 2026. Different modality, different operator, so the two are not measuring the same thing, but the range is worth holding in mind: public leaderboards presented with identical confidence can sit two orders of magnitude apart on the evidence underneath them. At the other end, a video model leaderboard ranking 12 models on 1,394 votes and OpenArt's arena, which publishes no vote counts at all, make 11,311 look generous.

The useful takeaway is not that this board is wrong. It discloses its rating formula, its voting interface and its total sample, which puts it ahead of most. The takeaway is that a three-digit arena score built on a four-figure vote count and an uncertainty penalty is a provisional signal, and anyone quoting 506 versus 362 as a settled quality verdict is reading more into it than the published method supports.

Sources: LLM Stats, Arena.


PS

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.