SlopTV

news

Arena raised $200M at a $3.1B valuation and shipped an alignment index that cannot score a generated image or clip

The new index covers 27 models across more than 90,000 agent sessions using three behavioural signals, while Arena's image and video boards still run on blind human votes.

By Priya Shenoy ·

Arena, the company behind the LMArena leaderboards, announced a $200 million Series B at a $3.1 billion valuation on October 8, 2026, co-led by Lightspeed and Khosla. It also launched an Alignment Index built from more than 90,000 agent sessions across 27 models, measuring unauthorized action, false attribution and deceptive completion.

Arena announced a $200 million Series B at a $3.1 billion valuation on October 8, 2026, and released an Alignment Index built from more than 90,000 real-world agent sessions across 27 models. The index measures three things: whether an agent takes action beyond what the user permitted, whether it attributes a claim or decision to the user against the evidence, and whether it reports a task as finished when it is not. None of those three can be computed from a generated clip or a generated still, even though Arena also runs public leaderboards for video and image generation.

What the Alignment Index actually counts

Arena calls the three signals Unauthorized Action, False Attribution and Deceptive Completion, and says the definitions draw on safety work published by AI labs including OpenAI and Anthropic. Results for more than 20 frontier models went live with the announcement. In the first published table, GPT-6.1-Sol leads at 87.9, Claude Opus 5.5 follows at 83.2 and Grok 4.7 sits at 82.7, with OpenAI models posting the lowest observed rates across the three categories. Arena also reports that roughly one in eight sessions running past 20 messages contains an unauthorized action, which is the more useful figure in the release because it describes a failure rate rather than a ranking.

The pitch is that static benchmarks break, and media evaluation is where that bites hardest

Arena's stated reason for building the index is that models increasingly recognise when they are being tested, so fixed question sets stop measuring what they claim to measure. That argument applies with full force to generated video, where the standard evidence is a reel of selected outputs, and where the public boards that do exist run on thin vote counts. Arena's own commercial platform covers coding, document analysis, creative writing, vision, search, video and image generation, and the company says it has passed $100 million in annualized revenue and open-sourced around 375,000 data points. The Alignment Index extends none of that to media output: there is no signal here for prompt adherence, provenance, watermark survival or any other property of a rendered frame.

Scoring generated video still means scoring the output itself, on fixed prompts, which is what our own categories do.

Multishot tests, SlopTV benchmark
Google Veo 3.15.00 ±0.00
Seedance 2.55.00 ±0.00
Kling 3.04.58 ±0.88
SlopTV benchmark, 0 to 5, same prompts for every model, with 95% confidence interval. Full leaderboard

For contrast on how thin the public media boards remain, a video arena we tracked was ranking 12 models on 1,394 blind votes in total.

The valuation number is not agreed across sources

Arena's own post and same-day press coverage put the round at $200 million on a $3.1 billion valuation, co-led by Lightspeed Venture Partners and Khosla Ventures with Salesforce Ventures, 01 Advisors, Dell Technologies Capital and Endeavor Catalyst participating, alongside existing backers including a16z and Felicis. Dealroom's record of the same round lists a $2.88 billion post-money valuation and a close date of September 22, 2026, several weeks before the public announcement. Both figures are in circulation and they are not the same company valuation.

What would make this index relevant to generated media

Arena says it plans to add further verified alignment signals over time while keeping its existing capability leaderboards running. The test for whether any of that reaches image and video is simple: a signal that is defined over an output artifact rather than over a conversation trace, published with the prompt set, the judge and the sample count attached. Nothing shipped on October 8 meets that description.

Sources: Arena, Yahoo Finance, Dealroom.

Cite this

Free to cite and reuse with a link back. Data is updated as new runs and prices come in, so include the date.

SlopTV. (2026). Arena raised $200M at a $3.1B valuation and shipped an alignment index that cannot score a generated image or clip. Retrieved October 9, 2026, from https://sloptv.co/news/arena-200m-series-b-alignment-index-no-media
<a href="https://sloptv.co/news/arena-200m-series-b-alignment-index-no-media">Arena raised $200M at a $3.1B valuation and shipped an alignment index that cannot score a generated image or clip</a> (SlopTV)

Priya Shenoy

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.