news
Comet's Opik adds a judge that returns a probability and cannot read images
Jev answers yes/no questions about text with a calibrated number instead of written reasoning, and Opik's own model guidance still points elsewhere for traces containing images.
Comet has wired a judge into Opik that never writes a sentence of critique, and the same documentation page that explains how to use it tells you to pick something else when your traces contain images. That second part is the one that matters for anyone trying to score generated pictures or video frames, because the cheap, fast, auditable judging path Comet is building runs on text only.
Jev answers yes or no, and shows its confidence as a number
The model is Jev, described in Opik's documentation as TypeSafe AI's System One classifier, served through OpenRouter. Comet's docs are blunt about what it does: Jev does not write free-form text, it answers yes/no questions about a piece of text and returns a calibrated probability, which Comet positions as a fast, low-cost judge for pass/fail checks. In the scoring pipeline, a probability at or above 0.5 maps to a 1, and the probability itself is preserved in the reason field rather than discarded.
That design removes most of the knobs a conventional LLM judge exposes. Jev is a single-shot classifier with no chat, no tool calling and no structured output, so Opik rejects configurations that assume those features, and it hides temperature, seed and the other model settings outright because they do nothing. Switching an existing rule over to Jev drops any non-Boolean scores and merges multiple text messages into a single user message, with a notice listing every change the switch made.
The prompt carries the data, the score description carries the question
The most counterintuitive part of the integration is where the question lives. The prompt template holds nothing but the payload, in the form of INPUT and OUTPUT placeholders, while each score's description supplies the actual question being asked, such as whether a reply addresses the user's question. Put the question into the prompt instead and Jev treats it as part of the text under evaluation rather than as an instruction. Every score in a rule is answered in a single call.
There is a hard ceiling of 32k on what Jev will read, and an evaluation that exceeds it is skipped with the reason recorded in the rule logs. Support shipped ahead of its documentation, routing the typesafe/jev-latest and typesafe/jev-1.13 identifiers to OpenRouter's Decisions API, which serves these models on a separate endpoint and rejects them on chat/completions. A feature toggle that would have gated the capability was removed before the code merged, which means Jev is available by default rather than behind a flag.
Why the vision gap bounds this for generated media
Opik's guidance on choosing a judge model notes that evaluating traces with images requires a model with vision capabilities, and Jev is not one. The split is becoming a pattern across this class of tool: Cloudflare's Clef decision models do accept images, while an open-weights scorer from AWS reads text only. Jev falls on the text side, so an image or video pipeline can use it on captions, prompts, metadata and tool arguments, but not on the frames themselves.
Comet's own side-by-side ran Jev and a gpt-4o-mini judge as online rules at 100% sampling over the same production traces, with one Boolean score and the same question put to both, and reports Jev as roughly 3.5 times cheaper and 3.8 times faster. That is the vendor measuring a model it has just added to its own product, on its own traffic, with the comparison baseline of its choosing, so the ratio is a starting point rather than a settled result. Comet frames judging as following embeddings, reranking and moderation into small dedicated models, and calls Jev the first of its kind that Opik supports directly.
Sources: Comet, Opik documentation, Opik on GitHub.
Cite this
Free to cite and reuse with a link back. Data is updated as new runs and prices come in, so include the date.
SlopTV. (2026). Comet's Opik adds a judge that returns a probability and cannot read images. Retrieved October 6, 2026, from https://sloptv.co/news/comet-opik-jev-classifier-judge-no-vision<a href="https://sloptv.co/news/comet-opik-jev-classifier-judge-no-vision">Comet's Opik adds a judge that returns a probability and cannot read images</a> (SlopTV)Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.