SlopTV

news

AWS open-sources a 2B model that scores options with a probability, and it cannot read images

Strands Decider 2B ships Apache 2.0 weights, training data and training scripts, but its inputs are text state and typed questions, which rules it out as a frame-level judge.

By Priya Shenoy · · Updated

AWS Strands Labs released Strands Decider 2B on October 1, 2026. It is a 1.9 billion parameter open-weight model that generates no text at all, instead picking one option from a supplied list and returning a calibrated probability. Weights, training data and training scripts are public under Apache 2.0. Inputs are text only.

AWS Strands Labs published Strands Decider 2B on October 1, 2026, a model whose entire design goal is to not write anything. It reads a block of state plus a typed question, picks one of the options handed to it in the request, and returns that choice with a probability attached. For anyone building automated scoring of generated images or video, the detail that decides whether this is useful is in what the announcement does not describe: there is no image or video input, only text state and typed questions.

The language head is removed and replaced with a roughly 1M parameter pointer head

The team starts from a Qwen3.5-2B-Base torso and discards the language modelling head, the component that turns hidden states into next-word predictions. In its place sits a pointer head of just over a million parameters, which scores the hidden state at each option position against the hidden state at the answer position. The torso carries a rank-16 LoRA adapter and the head runs in fp32. One forward pass produces the result, with no decoding loop, and the label set comes from the request rather than from the model, so there is no fixed cap on how many options you can offer. The released checkpoint is v19, and the team has said an earlier slot-head design performed noticeably worse.

Open here means weights, data and scripts, not a hosted endpoint

The release is Apache 2.0, with weights on Hugging Face and the code, training data and training scripts in the strands-labs repository on GitHub. Installing the package gives a command line tool and an HTTP server, and that server binds to localhost with no authentication, so running it as anything other than a local process means adding your own auth layer. No hosted inference provider is serving it, which means the cost model is your hardware rather than a per-call price.

The latency figures in circulation do not agree with each other

Amazon's own post says the model can answer meaningful questions in tens of milliseconds. Secondary coverage of the same release lands higher, around 106 to 115 milliseconds median on an Nvidia RTX 3090, with roughly 153 milliseconds reported for small tasks on an M3 MacBook. All of those are vendor-run or vendor-derived numbers on unstated prompt sizes, and none of them is an independent measurement, so the honest reading is a range rather than a figure.

The accuracy claim in the announcement is comparative, not numeric

The post states that the model's accuracy and calibration are competitive with the other models it knows of in this class, and leaves it there. No accuracy table, no baseline list, no test set named in the announcement itself. A JevBench score of 0.723, or 167 correct out of 231 tasks, has circulated in secondary write-ups, but it does not appear in the primary post, which is a strange omission for a release that otherwise ships its evaluation harness in the open.

Why a text-only chooser still matters to anyone scoring generated media

The architecture is close to what a practical media judge should look like. A closed label set means the output is always parseable, a probability per label means you can threshold and audit it, and a sub-second local answer means you could afford to run it on every clip in a batch rather than sampling. The constraint is the input. Strands Decider 2B can score a caption, a prompt, a moderation record or a metadata blob, but it cannot look at a frame, which leaves it a layer away from the work. Cloudflare's Clef, released the same day, takes images and video alongside text, and that single difference is what separates a routing component from a candidate evaluator. If the Strands team adds a vision encoder to this torso, the open training scripts mean the rest of the recipe is already on the table.

Sources: Strands Agents, MarkTechPost.

Cite this

Free to cite and reuse with a link back. Data is updated as new runs and prices come in, so include the date.

SlopTV. (2026). AWS open-sources a 2B model that scores options with a probability, and it cannot read images. Retrieved October 3, 2026, from https://sloptv.co/news/aws-strands-decider-2b-open-weights-text-only-judge
<a href="https://sloptv.co/news/aws-strands-decider-2b-open-weights-text-only-judge">AWS open-sources a 2B model that scores options with a probability, and it cannot read images</a> (SlopTV)

PS

Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.