SlopTV

news

An image-understanding leaderboard dated October 1 carries sourced scores for only eight models

NVIDIA's Nemotron 3 Nano Omni 30B A3B leads BenchLM's image-understanding slice at 84.6, and seven of the eight ranked models are open weights.

By Priya Shenoy ยท

BenchLM's image-understanding ranking, dated October 1, 2026, lists eight models with sourced benchmark scores spread across 13 benchmarks. NVIDIA's Nemotron 3 Nano Omni 30B A3B leads at 84.6, ahead of Qwen3.8 Max at 83.5 and Qwen3.6-27B at 81.7. Seven of the eight are open weights, and the site itself calls the coverage sparse.

Any automated judge of generated images rests on a vision stack that has to read the frame before it can grade it, and the public benchmark evidence for that first step is thinner than the volume of leaderboard coverage suggests. BenchLM's image-understanding ranking, dated October 1, 2026, lists eight models with sourced benchmark scores drawn from a reporting family of 13 benchmarks, and the page's own summary says the coverage is still sparse.

Eight entries, a 2.9-point gap at the top, and an open-weight sweep

NVIDIA's Nemotron 3 Nano Omni 30B A3B leads with a sourced average of 84.6, followed by Alibaba's Qwen3.8 Max at 83.5 and Qwen3.6-27B at 81.7. Qwen3.6-35B-A3B sits fourth at 80.6 and Qwen3.7 Plus fifth at 76. The lower half drops off sharply: Zyphra's ZAYA1-VL-8B at 74.8, then LiquidAI's LFM2.5-VL-3B at 57.7 and LFM2.5-VL-450M at 55.7, a spread of nearly 29 points between first and last.

Two things stand out in that table. Seven of the eight entries are open weights, with Qwen3.7 Plus the only proprietary model carrying a sourced score, and four of the eight come from a single family at Alibaba. The leaders are also small by frontier standards, with the top model at 30B total parameters and 3B active. Nothing from OpenAI, Google or Anthropic appears with a sourced score in this slice at all, which is less a statement about those models than about what gets published in a form a tracker can cite.

Estimated scores and a 12% multimodal weight on the main board

The same site's composite ranking, BenchAlign, shows how much of the multimodal picture is inference rather than measurement. Multimodal counts for 12% of the composite and is built from MMMU-Pro, AA-MMMU-Pro, OfficeQA Pro and CharXiv, four weighted benchmarks alongside 65 context-only ones. Claude Opus 5.5 appears near the top at 87.78 with a multimodal subscore of 89, but the entry is flagged as estimated and carries a conditional range of 73.4 to 100.0. A range that wide is an honest disclosure and a warning at the same time, since a model placed anywhere inside it would land in a completely different tier.

Reading an image and scoring a generated one are different jobs

These benchmarks test document, chart and diagram comprehension rather than aesthetic or fidelity judgment of synthetic output, so none of them directly answers whether a given model is a good grader of a generated picture. They are still the closest public proxy for the perception layer that any image judge inherits. A scoring harness built on a model that misreads what is in the frame will produce confident numbers about the wrong content, and that error never shows up in the final rubric score.

The practical takeaway for anyone picking an evaluator is that benchmark averages and preference data answer different questions. Arena-style boards for generated images rest on human votes, with one such board running on 11,311 blind votes across 10 models, while this ranking is a sourced average with no human comparison in it. Meanwhile the shift toward purpose-built scoring models that read images, rather than general chat models prompted into the role, continues; Cloudflare's Clef returns probabilities and accepts image input. Neither approach is covered by an eight-model table, and the table says as much about itself.

Sources: BenchLM, BenchLM leaderboard.

Cite this

Free to cite and reuse with a link back. Data is updated as new runs and prices come in, so include the date.

SlopTV. (2026). An image-understanding leaderboard dated October 1 carries sourced scores for only eight models. Retrieved October 4, 2026, from https://sloptv.co/news/benchlm-image-understanding-eight-models-sourced
<a href="https://sloptv.co/news/benchlm-image-understanding-eight-models-sourced">An image-understanding leaderboard dated October 1 carries sourced scores for only eight models</a> (SlopTV)

PS

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.