█

AI/TLDR

Strands Decider

AWS's small open decision model that picks, answers yes/no or rates with a calibrated confidence, locally

Structured OutputOpen source
Latest
2B hobson-v19
Updated
1 Oct 2026
Language
Python
License
Apache-2.0
Coverage
1 story
$pip install strands-decider

What's new

2B hobson-v191 Oct 2026

AWS Strands Labs released Strands Decider 2B: an Apache-2.0 decision model on Qwen3.5-2B with a pointer head that answers choice, yes/no and score questions with calibrated confidence in about 115 ms on an RTX 3090.

Latest news

Overview

Strands Decider is a "system one" decision model for agentic workflows from AWS Strands Labs, the experimental arm of the Strands Agents project. It does not generate text. You give it a piece of state and one or more questions — pick one of these options (`choice`), yes or no (`noul`), or rate on an ordered scale (`score`) — and it returns an answer with a calibrated confidence for every option.

Strands Decider 2B architecture: a Qwen3.5-2B torso with the LM head replaced by a pointer head
Strands Decider 2B model architectureStrands Agents blog ↗

The reference model, `strands-decider-2B-hobson-v19`, starts from the Qwen3.5-2B base model. The language-model head is removed and replaced with a pointer head of about one million parameters that scores each supplied option by comparing hidden states, plus a rank-16 LoRA adapter. One forward pass answers the question, with no decoding loop, so every output is one of your options and nothing has to be parsed.

The README reports 72.3% (167 of 231) on the public JevBench set with a Brier score of 0.342 and an ECE of 0.052, a 115 ms median latency per question on an RTX 3090 and a 153 ms warm median on an M3 Pro Mac. The code, training setup and data source list are Apache-2.0, and a full training run takes about 11 hours on one RTX 3090. Within structured output it sits beside Jeff, Laya and Jeeves, which answer the same Jev-style questions.

Chart of Strands Decider decision latency rising with task size in tokens
Decision latency as a function of task sizeStrands Agents blog ↗

What it does

  • Three question types — `choice`, `noul` (yes/no) and `score` — and several can be asked about the same state in one call
  • A calibrated confidence on every answer; the README reports about 95% accuracy on unseen tasks when confidence is 0.9 or higher
  • Single forward pass through Qwen3.5-2B with a pointer head, so outputs are always one of the supplied options
  • `strands-decider ask` CLI and a `strands-decider serve` HTTP server with a `/v1/systemone` endpoint
  • Runs on NVIDIA GPUs, on Apple silicon with optional MLX acceleration, and on CPU
  • Training scripts included: about 11 hours on one 24 GB RTX 3090 or about 1 hour 10 minutes on eight H100s

Getting started

Strands Decider installs from PyPI and downloads the reference model from Hugging Face on first use.

Install the package

bashbash
pip install strands-decider

Ask a choice question

The answer comes back with a probability for each option.

bashbash
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days!" \
  --choice "Which team should handle this?=billing,sales,retail"

Run it as a server

bashbash
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000

Send a yes/no question over HTTP

bashbash
curl -s localhost:8000/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "Help! My payouts have been failing for 3 days!",
    "questions": {
      "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
    }
  }'

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Routing a request to the right model or team before a larger LLM is called
  • Choosing a tool and checking its arguments before an agent runs it
  • Guardrails and policy classification with a confidence threshold
  • Scoring whether an agent's output is good enough, as a cheap local evaluator

How Strands Decider compares

Strands Decider alongside other open-source structured output tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Laya★ 30.4kAn Apache-2.0 decision model that answers choice, score and boolean questions about text in one forward pass, returning calibrated probabilities instead of generated JSON.
Guidance★ 21.8kA programming model that interleaves generation, prompting, and control logic to constrain output and enforce formats like JSON or regex patterns.
Outlines★ 15.9kA library for structured generation that constrains an LLM's token output to match a JSON schema, regex, or grammar so the result is always valid.
Instructor★ 14kA library that wraps an LLM client to return data validated against a schema, retrying automatically on invalid output, with SDKs in several languages.
BAML★ 9.4kA domain-specific language for defining LLM functions with typed schemas, parsing flexible model output into reliable structured data across many languages.
Marvin★ 6.2kA Python toolkit from Prefect for turning LLM calls into typed functions that extract, classify, and cast text into structured Python objects.
SemIf★ 4.7kA decision baseline that reads typed option probabilities directly from an open model in one forward pass, instead of generating and parsing JSON.
Strands Decider—AWS's small open decision model that picks, answers yes/no or rates with a calibrated confidence, locally