A 2B model that only picks from options

Strands Decider 2B answers typed questions about a piece of text. It never writes text back. AWS's Strands Labs introduced it on 1 October 2026 as a decision model.

The repository lists three question types:

  • noul, a yes or no question that returns a value between 0 and 1
  • choice, which picks one of N options with a confidence
  • score, which places the text on an ordered rubric

The model takes the Qwen3.5-2B-Base torso, drops its language model head and adds a pointer head of about a million parameters. The torso gets a rank-16 LoRA adapter. The repo puts the total at 1.9 billion parameters.

Every answer comes out of one forward pass. Several questions about one text are cheap, the README says, because the text is read once.

Install it with pip install strands-decider. The CLI has ask for one-off questions and serve for a local HTTP endpoint at /v1/systemone.

Decision models arrived in September

The Strands post says decision models have drawn attention "since TypeSafe AI's launch of Jev earlier this month." For how that model drives a browser loop, see Jev Ultrafast: Browser Use wires TypeSafe Jev into an indexed browser-action loop.

OpenAI put a hosted Decisions API into limited preview at DevDay on 29 September US time. See OpenAI DevDay 2026 brings Ultrafast, a Decisions API preview, Codex cloud and plugin extensions.

Strands Decider is the open, local option in that group. The team says it has seen early use for model routing, tool selection, guardrails, evaluations and policy classification.

Apache 2.0, six commits, usable for local experiments

The code and the Hugging Face checkpoint both carry the Apache License 2.0. The repo lists its training data sources and their licences, plus the full training recipe.

When The Frontier checked at 15:48 SGT on 3 October, the repo had 232 stars and six commits. PyPI had version 0.1.0, uploaded 2 October SGT, and calls the package "an experimental typed classifier."

By our reading, it is usable for local experiments and not yet production-shaped. The README says the HTTP server binds to 127.0.0.1 and has no authentication.

The repo reports these v19 figures:

Measurev19 result
JevBench v1 public set accuracy0.723, or 167 of 231 tasks
Brier score and expected calibration error0.342 and 0.052
Easy, standard and hard tiers1.000, 0.875 and 0.505
Median and p95 latency, RTX 3090115 ms and 299 ms
Warm median latency, M3 Pro153 ms under 300 tokens

The team says retraining the full recipe takes about 11 hours on one RTX 3090.

Pick it for cheap gates around an LLM

The README says frontier LLM APIs expose nothing like its calibrated confidence. On short classification tasks it had not seen, it says answers at 0.9 confidence or more were right about 95% of the time.

Pick it when an agent needs a fast yes or no before a step runs. The repo's worked example gates a weather tool call on two questions. Are the arguments grounded in what the user said? Is it too early to call the tool?

Stay with an LLM for anything that needs generated text, code or multi-step reasoning. The Strands post says the single parallel pass makes the model "significantly worse at solving complex problems than reasoning models."

A hosted service like OpenAI's preview may suit teams that do not want to run a GPU. Strands Decider suits teams that want local control and can retrain.

Hard tasks and small samples

The hard tier score of 0.505 means the model got about half of the hardest public JevBench tasks right.

The team warns that 231 tasks are few. Six retrains of an earlier recipe varied by a standard deviation of 3.2 tasks, so it treats gaps under about 10 tasks as unresolved.

The reference model is v19. A later run, v20, scored 169 against 168 at a 4096-token window and failed four of its preregistered predictions, so it did not replace v19. Pin the checkpoint name when you test.

All performance figures here come from the Strands team, and The Frontier found no independent test. That includes the post's placement of third among 33 models in the 2B class.

The MLX extra for Apple silicon ships in the next release. Until then, the README says to install it from a clone.

Strands post, code, weights and package