An AI agent deciding which tool to call next does not need to write an essay. A growing class of models now skips text generation entirely: they read a situation, pick one answer from a fixed list, and report how confident they are. AWS’s Strands Decider 2B, released on October 1, 2026, is the clearest sign yet that “decision models” are becoming a standard layer inside AI agents — and this explainer breaks down how they work, what the published numbers really mean, and when you should use one instead of a large language model.
What is a decision model?
A decision model is a small AI model that chooses between predefined options, or rates something on a scale, instead of generating text. The category took shape after TypeSafe AI launched its model Jev in September 2026, and the field also calls them “System One models,” according to MarkTechPost’s October 1, 2026 technical write-up of the Strands release. Every answer comes from the allowed set of options and carries a calibrated confidence score, so calling code can decide whether to accept the choice or escalate it. The model never writes a paragraph, which is precisely the point: for routing, triage and guardrail steps, fluent text is wasted work.
How does Strands Decider 2B work?
Strands Decider 2B takes an open 2-billion-parameter language model and removes the part that produces words. The Strands Labs team started from Qwen3.5-2B-Base, discarded the language-modelling head, and replaced it with a pointer head of roughly 1 million parameters that compares the model’s internal state against each candidate answer, per MarkTechPost (October 1, 2026). One forward pass returns the result — there is no token-by-token decoding loop, which is why it is so fast. The model supports three question types: choice (pick 1 of N options), noul (a yes/no probability between 0 and 1), and score (a level on an ordered rubric). The weights, training data list and training recipe are published on Hugging Face (checkpoint StrandsAgents/strands-decider-2B-hobson-v19) under an Apache-2.0 licence, and pip install strands-decider provides a command-line tool and a local HTTP server. One practical caveat from MarkTechPost’s write-up: the bundled server binds to localhost with no authentication, so any shared deployment needs your own access controls.
“What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — ‘what is the next thing for me to do here, based on where I am?'” — Marc Brooker, AWS distinguished engineer, speaking to TechCrunch, October 1, 2026
Brooker, who built the first version after trying TypeSafe’s Jev, told TechCrunch the demand came from AWS customers whose agent workflows did not need “the capability or cost of a fully featured LLM all the time.” If your agents spend much of their budget on access control and tool routing, our AI agent permissions guide covers the safety side of the same problem.
How fast and accurate is Strands Decider 2B?
On the public JevBench set, the released v19 checkpoint scores 0.723 accuracy with a median latency of 115 milliseconds on a consumer RTX 3090 GPU, according to figures published by the Strands team and reported by MarkTechPost on October 1, 2026. The calibration numbers matter more than raw accuracy for this class of model: the Brier score is 0.342 and the expected calibration error is 0.052, and the team reports that answers given at 0.9 confidence or higher were right about 95% of the time on unseen short classification tasks. Accuracy by difficulty tier was 1.000 on easy tasks, 0.875 on standard tasks and 0.505 on hard tasks — a number we return to below. On the September 25, 2026 JevBench board (v1.4.2), v19 ranked 3rd of 33 models in the 2B class. Two honesty notes: these are vendor-published benchmark figures that OpenAIMaster.ai has not independently reproduced, and JevBench is a third-party benchmark built for this model category, so treat the ranking as indicative rather than definitive.
How do decision models compare?
Decision models trade the open-ended ability of an LLM for speed, predictable output shape and a usable confidence score. The table below compares the open decision models in MarkTechPost’s October 1, 2026 roundup; latencies come from different hardware and harnesses and are not directly comparable.
| Model | Developer | Access | Size | JevBench accuracy |
|---|---|---|---|---|
| Strands Decider 2B (v19) | Strands Agents (AWS) | Open weights, Apache-2.0 | 1.9B | 0.723 |
| Jev 1.13.0 | TypeSafe AI | Closed hosted API | Undisclosed | Not in source table |
| decider-2b | Mapika | Open code/weights | 1.9B | 0.710 |
| Decision 2B | FlyMy.AI | Open code/weights | 2.5B | 0.753 |
Against a frontier LLM, the trade is starker: an LLM can answer questions nobody pre-listed, but each call costs more, returns free-form text that needs parsing, and gives you no reliable “how sure are you” signal. A decider can only answer questions you framed in advance — and that constraint is its reliability feature, as Brooker put it to TechCrunch: “more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and] lower latency, potentially lower cost.” For a broader view of this month’s agent tooling, see our roundup of AI tools for October 2026.
When should you use a decider instead of an LLM?
Use a decision model when the answer is one of a known set, the step repeats many times per task, and you can act on a confidence score. That is our three-question test, built from the use cases the Strands team reports (model routing, tool selection, argument checking, triage, guardrails and evaluations) and from the published calibration data:
- Is the answer one of a known set? Routing a support ticket to billing, sales or retail; picking which tool an agent should call; checking whether an action passes a policy. If yes, a closed domain fits. If the step needs writing, coding or summarising, it does not — the Strands team itself states the model is unsuited for those jobs.
- Does the step repeat dozens of times per task? A single agent run can make many routing calls. At a 115 ms median on local hardware (and 153 ms warm on an Apple M3 Pro, per the team’s figures), a decider removes both the latency and the per-call API cost of asking a frontier model “what next?” over and over.
- Can you act on the confidence score? The practical pattern is a hybrid agent: accept decider answers at high confidence, and escalate the rest to a bigger model. The team’s own guidance is to confirm or escalate below roughly 0.9 confidence, where their measurements show accuracy dropping off.
If all three answers are yes, a decider is very likely the cheaper, faster, more predictable choice today. One caveat we want to be explicit about: OpenAIMaster.ai has not hands-on tested Strands Decider 2B yet, so this framework rests on the published architecture and benchmark reporting cited throughout, not on our own benchmark run.
Why is this happening now?
Decision models are hardening into an infrastructure tier because agent builders are hitting the cost of using frontier models as glue. AWS did not move alone: Cloudflare shipped its own decision-model pair, Clef and Clef-flash, on the same day, and OpenAI announced a similar offering the same week, according to TechTimes (October 2, 2026) and TechCrunch (October 1, 2026). TypeSafe named Jev after the economist William Stanley Jevons, invoking his theory that when the cost of something falls — here, machine judgement — demand for it rises. Not everyone is convinced the clones match the original: TypeSafe CEO Diogo Almeida told TechCrunch, “I get that people think it’s a gold rush, but they might be underestimating the difficulty of making the models actually smart.” The governance stakes of giving agents faster, cheaper decision-making are real, too — see our explainer on why regulators are ending the “AI did it” defence.
What are the limits?
A decision model is a component, not a brain, and its published hard-task accuracy shows why the escalation path matters. The 0.505 accuracy on JevBench’s hard tier is barely better than chance, so a decider should never be the final word on a difficult or high-stakes call. It cannot write code, hold a conversation or summarise a document. There is no hosted inference endpoint yet, so you run it yourself — including adding authentication the bundled server lacks. And the headline benchmark is the vendor’s own JevBench run: credible as a starting point, not yet independent verification.
Frequently asked questions
What is a decision model in AI?
A decision model is a small model that picks between predefined options, answers yes/no with a probability, or scores on a rubric — always with a confidence value — instead of generating text. Decision models are also called System One models, a category named after TypeSafe AI’s Jev launched in September 2026.
Is Strands Decider 2B free to use?
Yes. AWS’s Strands Labs published the weights, training data list and scripts under the Apache-2.0 licence on Hugging Face, so commercial use, modification and redistribution carry no licence fee. You do run it on your own hardware or cloud instances.
Can a decision model replace an LLM like ChatGPT or Claude?
No. The Strands team states Decider 2B is worse than reasoning models on complex problems and unsuited to coding, chat or summarisation. The intended design is a hybrid agent: the decider handles fast, repetitive choices and escalates hard or open-ended work to a large language model.
How accurate is Strands Decider 2B?
The published v19 checkpoint scores 0.723 accuracy on the public JevBench set (167 of 231 tasks), with accuracy of 1.000 on easy, 0.875 on standard and 0.505 on hard tasks, per figures reported by MarkTechPost on October 1, 2026. These are vendor-published results, not independently verified.
Sources and methodology
This explainer rests on the Strands team’s published architecture and JevBench figures as reported by MarkTechPost and TechCrunch on October 1, 2026, and on TechTimes’ October 2, 2026 report for the Cloudflare launch; claims were verified against those sources on October 4, 2026. Benchmark figures are vendor-reported and labelled as such; OpenAIMaster.ai has not independently run JevBench or tested the model hands-on.
- TechCrunch — Amazon releases its own Jev clone as decision models flood the web (October 1, 2026)
- MarkTechPost — AWS Strands Labs Releases Strands Decider 2B (October 1, 2026)
- TechTimes — AWS Releases Decision Model for AI Agents That Routes Without Generating Any Text (October 2, 2026)
OpenAIMaster is an independent publication covering artificial intelligence — from model launches and AI news to hands-on tool reviews and practical guides. Our testing methodology is published openly, and every review is updated as tools evolve.
Feel free to email us at contact@openaimaster.ai — we are happy to help!


