SlopTV

news

Cloudflare's Clef returns probabilities instead of prose, and it reads images

The two open-weight decision models released October 1 answer a fixed schema of typed questions with a score per allowed answer, but the only published timing comparison is a website classification task, not anything to do with generated media.

By Priya Shenoy · · Updated

Cloudflare released Clef and Clef-flash on October 1, 2026, open-weight decision models hosted on Workers AI under Apache 2.0. Instead of writing text, they take an input and a schema of typed questions and return a probability for every allowed answer. Both accept image input, and Cloudflare claims the top spot on the Jev Decision Index.

Scoring a generated image with a language model usually means asking it for prose or JSON, parsing whatever comes back, and trusting that the number it wrote down means something stable. Cloudflare released two models on October 1 that remove the writing step. Clef and Clef-flash take a piece of state plus a schema of typed questions and return a probability for every allowed answer, and both of them accept image input.

Two models on Workers AI, Apache 2.0 weights, and an RL fine-tuning service

Cloudflare describes Clef and Clef-flash as Cloudflare-trained decision models hosted on Workers AI, released on Hugging Face under an Apache 2.0 license so they can be run locally, and compatible with the Jev API so existing integrations can be repointed without a rewrite. A reinforcement learning fine-tuning platform shipped alongside them, aimed at developers who want to tune a decision model on their own labelled data rather than prompt a general model into behaving like a classifier.

The one timing number Cloudflare published is about website categories

The concrete measurement in the announcement comes from Cloudflare's own Threat Intelligence team, which has been using Clef to classify domains. Given a domain, the model returns a spread such as 95 percent fashion, 85 percent ecommerce and under 1 percent phishing. Cloudflare puts that workflow at 2.2 seconds end to end including fetching and rendering the page, against 4.7 seconds for gpt-oss-120b, which returned only two classifications in the same run. That is a single internal task with no sample count attached, so it establishes a speed gap on one workload rather than general accuracy.

The leaderboard claim points at a demo site, not a table

Cloudflare states that Clef currently leads the Jev Decision Index and directs readers to a live benchmark demo site for full results, rather than printing per-task scores, sample counts or confidence intervals in the post itself. The index shares its name with the product family whose API Clef copies, which is worth keeping in mind when reading a first-party claim of the top position. Parameter counts, context length and median latency figures for both models circulated widely in same-day writeups, and those figures are not consistent with each other across sources, so the demo site and the model cards are the only places to settle them. The training data and pipeline were not published, making this an open-weights release rather than a reproducible one.

Nothing here has been tested on generated video or images

For anyone scoring synthetic media, the appeal is obvious. A rubric like prompt adherence, visible artifacts, hand and face integrity, or text legibility is a bounded set of typed questions, which is exactly what this model shape consumes, and a calibrated probability per answer is more useful than a judge's self-reported 7 out of 10. The gap is evidence. Cloudflare published no result on image quality, no artifact detection task, and no agreement rate against human raters, and the vision encoder comes from the underlying backbone rather than from any media-specific training. A judge that has not been checked against human preference data on generated frames is a classifier with a plausible interface, not a validated evaluator, which is the same objection that applies to any scoring system that ships without its prompt set and vote counts, including several creative model leaderboards running today.

The useful next step is cheap: take an existing set of human-voted image comparisons, such as the 11,311 votes behind a recent image arena refresh, and measure how often a decision model's probabilities agree with them. Until someone does that, Clef's relevance to media scoring is a hypothesis.

Sources: Cloudflare, Hugging Face.

Cite this

Free to cite and reuse with a link back. Data is updated as new runs and prices come in, so include the date.

SlopTV. (2026). Cloudflare's Clef returns probabilities instead of prose, and it reads images. Retrieved October 2, 2026, from https://sloptv.co/news/cloudflare-clef-decision-models-image-judging
<a href="https://sloptv.co/news/cloudflare-clef-decision-models-image-judging">Cloudflare's Clef returns probabilities instead of prose, and it reads images</a> (SlopTV)

PS

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.