The Model That Can't Write: AWS Open-Sources Strands Decider 2B, a 2B-Parameter Decision Engine for Agents
AWS's Strands Labs took a Qwen3.5-2B torso, deleted the LM head, and shipped a decision model that answers in tens of milliseconds on a laptop — fully open, weights, data, and training scripts included.
The most interesting model Amazon released this week cannot write a single word. Strands Decider 2B, open-sourced on October 1 by AWS’s Strands Labs, is a “decision model” — a deliberately crippled LLM that has been surgically stripped of its ability to generate text and rebuilt to do exactly one thing: pick the right answer from a list of options, attach a calibrated confidence score, and do it in tens of milliseconds on hardware you already own.
That may sound like a downgrade. It is actually the point. As AI agents have proliferated over the past two years, a growing share of their runtime work is not “write me a poem about quantum computing” but “is this tool call grounded in what the user actually said — yes or no?” Frontier LLMs are massively overqualified for those questions, and they charge accordingly, in latency and dollars alike. Strands Decider 2B is AWS’s bet that the future agent stack is hybrid: a big reasoning model for the hard calls, and a tiny decision model for everything else.
What a decision model actually is
The category is barely a month old in its current form. TypeSafe AI launched Jev earlier in September — named after the economist William Stanley Jevons, whose theory that falling cost increases demand the company borrows as a thesis — and dozens of imitators have shipped since. Cloudflare open-sourced Clef, its own homegrown entry, just a day before Amazon’s release. The shared insight is architectural: a language model’s next-token prediction head is wasted machinery when your output space is a closed set of options.
AWS’s implementation starts from a Qwen3.5-2B torso. The team removed the LM head entirely, taking away the model’s capacity to generate text, and replaced it with a “pointer head” that scores the hidden state at each candidate answer position against the hidden state at a dedicated <answer> position. The pointer head is tiny — just over a million parameters — and the torso itself is fine-tuned with a rank-16 LoRA adapter. Everything runs in a single parallel pass, which is where the speed comes from.
What the model gives back is tighter than a generation: given a state and a question, it returns a ranked distribution over allowed answers plus a reliability score. “Which team should handle this? — billing, sales, retail” comes back as billing (confidence 0.768), with the full score breakdown attached. Because the output space is closed, the model can never emit an invalid answer — a property that matters a great deal when the consumer is a workflow engine rather than a human.
The trade-off is real, and the Strands team is candid about it: generating all outputs in one parallel pass makes decision models significantly worse at complex multi-step problems than reasoning models, and their inability to produce text rules them out for coding, chat, summarization, and most classic LLM work. This is a screwdriver, not a Swiss Army knife.
The numbers: accuracy, calibration, latency
On JevBench’s public set (the model released today is v19 — the repo documents every architectural iteration, including an earlier “slot head” design the team found performed significantly worse), Strands Decider 2B ranks 3rd of 33 in the 2B class, and 1st of 30 when the just-over-2B models are excluded. Calibration, measured by Brier score on the same set, is competitive with anything else in the class — meaning the confidence scores are actually trustworthy, not decoration.
Latency is the headline. Median local decision time is around 115ms on an NVIDIA RTX 3090, and roughly 153ms for small tasks on an M3 MacBook with no discrete GPU. Latency scales approximately linearly with task size in tokens. For comparison, a round-trip to a frontier LLM API is rarely under several hundred milliseconds and often much worse; a decision this cheap can sit inline in a code path where an LLM call never could.
Why 2B parameters? Two reasons, per the team. First, experimentation: you can run and even train the thing on hardware you already have, which makes trying ideas fast and low-risk. Second, it’s a genuine sweet spot — small enough for local iteration, large enough to do meaningful work. The model scores 100% on JevBench’s “easy” task tier, and easy tasks map remarkably well onto the rote decisions that dominate agent workflows in production.
Born from a homebrew project
The origin story is unusually grassroots for AWS. Distinguished engineer Marc Brooker — a longtime distributed-systems voice inside Amazon — saw Jev, tried building his own take on the architecture, and the homebrew project was successful enough that it briefly held the top spot on the JevBench ranking for its size class. Amazon engineers then cleaned it up and released it through Strands Labs, the organization AWS spun up earlier this year to prototype new tools and protocols for deploying AI agents.
Brooker told TechCrunch the need came straight from customer conversations: agentic workflows that didn’t always justify the capability or the cost of a fully featured LLM. “What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — ‘what is the next thing for me to do here, based on where I am?’” The payoff is “a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, lower latency, potentially lower cost.”
The reference use case in the repo is telling. A deliberately over-eager demo agent, asked “What’s the weather?”, guesses a city and fires off the tool call anyway. Before the call executes, Strands Decider reads the conversation and the proposed call, and answers two yes/no questions: are these argument values grounded in anything the user actually said, and is it premature to call this tool now? A few lines of Python turn those predictions into a typed action — Proceed, Deny, Confirm (escalate to a human), or Guide (hand the turn back with feedback) — and the agent asks which city you meant instead of confidently reporting the weather somewhere nobody mentioned. That intervention hook is the existing Strands InterventionHandler pattern, and the same shape holds whether the decider is this model, a Cedar policy, or another agent entirely.
Open, actually
The release is open in the fullest sense: Apache-2.0 weights on Hugging Face, code on GitHub, and — unusually — all the training data and the scripts used to build the model. There’s a pip install strands-decider CLI, worked examples of integration inside a Strands agent, and the Strands team says libraries for decision-model integration are coming. Bloomberg-scale budgets are not required; Brooker notes the cost of building something interesting in this class runs to hundreds or thousands of dollars, which is why he doesn’t expect frontier labs to own the niche.
TypeSafe’s CEO Diogo Almeida, for his part, dismissed the current crop as “more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful” — while conceding, implicitly, that the gold rush is real.
The strategic read is straightforward. Every major agent framework is converging on the same problem: the expensive model should not be making the cheap decisions. Whether the decider that wins is Jev, Clef, Strands Decider, or something not yet built, the category it belongs to — small, fast, calibrated, closed-world — now has Amazon’s weight behind it, and a fully open reference implementation anyone can fork this afternoon.