Overview
Strands Decider is a "system one" decision model for agentic workflows from AWS Strands Labs, the experimental arm of the Strands Agents project. It does not generate text. You give it a piece of state and one or more questions — pick one of these options (`choice`), yes or no (`noul`), or rate on an ordered scale (`score`) — and it returns an answer with a calibrated confidence for every option.

The reference model, `strands-decider-2B-hobson-v19`, starts from the Qwen3.5-2B base model. The language-model head is removed and replaced with a pointer head of about one million parameters that scores each supplied option by comparing hidden states, plus a rank-16 LoRA adapter. One forward pass answers the question, with no decoding loop, so every output is one of your options and nothing has to be parsed.
The README reports 72.3% (167 of 231) on the public JevBench set with a Brier score of 0.342 and an ECE of 0.052, a 115 ms median latency per question on an RTX 3090 and a 153 ms warm median on an M3 Pro Mac. The code, training setup and data source list are Apache-2.0, and a full training run takes about 11 hours on one RTX 3090. Within structured output it sits beside Jeff, Laya and Jeeves, which answer the same Jev-style questions.

What it does
- Three question types — `choice`, `noul` (yes/no) and `score` — and several can be asked about the same state in one call
- A calibrated confidence on every answer; the README reports about 95% accuracy on unseen tasks when confidence is 0.9 or higher
- Single forward pass through Qwen3.5-2B with a pointer head, so outputs are always one of the supplied options
- `strands-decider ask` CLI and a `strands-decider serve` HTTP server with a `/v1/systemone` endpoint
- Runs on NVIDIA GPUs, on Apple silicon with optional MLX acceleration, and on CPU
- Training scripts included: about 11 hours on one 24 GB RTX 3090 or about 1 hour 10 minutes on eight H100s
Getting started
Strands Decider installs from PyPI and downloads the reference model from Hugging Face on first use.
Install the package
pip install strands-deciderAsk a choice question
The answer comes back with a probability for each option.
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
--state "Help! My payouts have been failing for 3 days!" \
--choice "Which team should handle this?=billing,sales,retail"Run it as a server
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000Send a yes/no question over HTTP
curl -s localhost:8000/v1/systemone \
-H 'content-type: application/json' \
-d '{
"state": "Help! My payouts have been failing for 3 days!",
"questions": {
"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
}
}'Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Routing a request to the right model or team before a larger LLM is called
- Choosing a tool and checking its arguments before an agent runs it
- Guardrails and policy classification with a confidence threshold
- Scoring whether an agent's output is good enough, as a cheap local evaluator
How Strands Decider compares
Strands Decider alongside other open-source structured output tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Laya | ★ 30.4k | An Apache-2.0 decision model that answers choice, score and boolean questions about text in one forward pass, returning calibrated probabilities instead of generated JSON. |
| Guidance | ★ 21.8k | A programming model that interleaves generation, prompting, and control logic to constrain output and enforce formats like JSON or regex patterns. |
| Outlines | ★ 15.9k | A library for structured generation that constrains an LLM's token output to match a JSON schema, regex, or grammar so the result is always valid. |
| Instructor | ★ 14k | A library that wraps an LLM client to return data validated against a schema, retrying automatically on invalid output, with SDKs in several languages. |
| BAML | ★ 9.4k | A domain-specific language for defining LLM functions with typed schemas, parsing flexible model output into reliable structured data across many languages. |
| Marvin | ★ 6.2k | A Python toolkit from Prefect for turning LLM calls into typed functions that extract, classify, and cast text into structured Python objects. |
| SemIf | ★ 4.7k | A decision baseline that reads typed option probabilities directly from an open model in one forward pass, instead of generating and parsing JSON. |
| Strands Decider | — | AWS's small open decision model that picks, answers yes/no or rates with a calibrated confidence, locally |