Overview
Strands Decider 2B is an open decision model, or "system one" model, published by AWS's Strands Labs on October 1, 2026, with the code on GitHub under Apache-2.0 and the weights on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19. Unlike an LLM it does not generate text: it picks between a set of options or rates something on a scale, and every answer carries a confidence. It is built for the decisions inside agentic workflows made with the Strands Agents SDK, such as model routing, tool selection, argument checking, triage, guardrails and evaluations.

The design keeps the torso of Qwen3.5-2B-Base, discards its language-modelling head and adds a pointer head of about a million parameters. The head scores each option by comparing the hidden state at the <answer> position with the hidden state at that option's last token, and a masked softmax turns the scores into per-option probabilities. The torso is adapted with a rank-16 LoRA adapter. Because the head has no per-option parameters, label sets come from the request rather than the weights and there is no cap on how many options a question may carry. Three question types share the same head: noul (yes/no), choice (one of N) and score (an ordered rubric).

The released checkpoint is v19. On JevBench's 231 public tasks it scores 0.723 (167 of 231), with a Brier score of 0.342 and an expected calibration error of 0.052, and 1.000 / 0.875 / 0.505 on the repository's easy, standard and hard tiers. The launch post places it 3rd of 33 in the 2B class on JevBench's public set. The README reports a median latency of 115 ms (299 ms at the 95th percentile) per JevBench question on an RTX 3090, and a 153 ms warm median on an M3 Pro Mac for prompts under 300 tokens. On short classification tasks it has never seen, answers at a confidence of 0.9 or more are right about 95% of the time.
It installs with pip install strands-decider and runs from a CLI or as a local server with a /v1/systemone endpoint, on CUDA, Apple silicon (MPS or MLX) or CPU. With --vision the server keeps Qwen3.5-2B-Base's vision tower so a request can carry images. The repository also holds the training recipe (about 11 hours on one RTX 3090, or 1 hour 10 minutes on eight H100s), the data sources and a preregistered record of every training run.
| Released | 2026-10-01 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 1.9B (Qwen3.5-2B-Base torso with a rank-16 LoRA adapter and a pointer head of about 1M parameters) |
| Architecture | Qwen3.5-2B-Base decoder torso with its language-modelling head discarded and replaced by a pointer head that scores each option against the hidden state at the <answer> position; one forward pass, no generation |
| Modalities | Text, Vision |
Benchmarks
Strands Decider 2B (v19) against the nearest open systems on JevBench's 231 public tasks, as published in the repository's evaluation notes. v19 is Strands Labs' own harness run; the others come from the benchmark's per-task data. Tiers are the repository's split of the public tasks.
| Benchmark | Strands Decider 2B (v19) | decider-2b | Open-Jev 2B | jeff 400M | Laya 421M | Open-Jev 9B |
|---|---|---|---|---|---|---|
| Overall (231 public tasks) | 0.723 | 0.71 | 0.645 | 0.628 | 0.584 | 0.775 |
| Easy tier | 1 | 1 | 1 | 1 | 0.958 | 1 |
| Standard tier | 0.875 | 0.847 | 0.764 | 0.75 | 0.694 | 0.903 |
| Hard tier | 0.505 | 0.495 | 0.414 | 0.387 | 0.351 | 0.595 |
| routing_hard | 1 | 1 | 0.8 | 0.6 | 0.2 | 1 |
| intent | 0.917 | 1 | 0.708 | 0.708 | 0.833 | 1 |
| adversarial | 0.667 | 0.667 | 0.5 | 1 | 0.667 | 1 |
| long_policy | 0.368 | 0.316 | 0.263 | 0.105 | 0.211 | 0.474 |
| multi_hop | 0.444 | 0.556 | 0.5 | 0.389 | 0.333 | 0.667 |
| policy | 0.75 | 0.833 | 0.917 | 0.917 | 0.75 | 1 |
| trap | 0.875 | 1 | 0.875 | 0.5 | 0 | 1 |
| judge_hard | 0.588 | 0.529 | 0.529 | 0.471 | 0.412 | 0.765 |
| temporal_numeric | 0.267 | 0.333 | 0.067 | 0.4 | 0.333 | 0.133 |
This model's scores
- JevBench v1, public set (231 tasks)72.3%
- JevBench, easy tier100%
- JevBench, standard tier87.5%
- JevBench, hard tier50.5%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- 0.723 on JevBench's 231 public tasks (167 of 231), with a Brier score of 0.342 and an expected calibration error of 0.052
- Every task in JevBench's easy tier answered correctly (1.000), and 0.875 on the standard tier
- Median 115 ms per JevBench question on an RTX 3090, and it also serves on an Apple-silicon Mac or on CPU
- Many questions about one text are cheap: the text is read once and each question adds only its own tokens
- Calibrated confidences: on unseen short classification tasks, answers at 0.9 confidence or more are right about 95% of the time
- Apache-2.0 code and weights, with the full training recipe retrainable in about 11 hours on one RTX 3090
Best for
- Reach for it to make the rote decisions inside a Strands agent, such as which tool to call next or whether a tool call's arguments are valid, and leave the hard decisions to an LLM.
- Reach for it to route an incoming request to the right team or queue, or to pick which LLM should handle a task, locally and in about a tenth of a second.
- Reach for it for guardrails and evaluations: grounding checks, policy classification and scoring model outputs at low cost, acting only on high-confidence answers.
- Look elsewhere for the hard tier of JevBench, where it scores 0.505 and the 9B Open-Jev scores 0.595, or for anything that needs generated text.
How to access
| Provider | Model ID |
|---|---|
| Hugging Face (weights) ↗ | StrandsAgents/strands-decider-2B-hobson-v19 |
| PyPI (strands-decider CLI and server) ↗ | — |
FAQ
What is Strands Decider 2B?
An Apache-2.0 decision model from AWS Strands Labs, released October 1, 2026. Instead of generating text, it picks between options (yes/no or one of N) or rates something on a scale, and returns a confidence with every decision. It is meant for the decisions inside agentic workflows built with the Strands Agents SDK.
How accurate is it?
The released v19 checkpoint scores 0.723 (167 of 231) on JevBench's public set, with a Brier score of 0.342 and an expected calibration error of 0.052. The launch post places it 3rd of 33 in the 2B class. The README notes that six retrains of an earlier recipe varied by 3.2 tasks, so differences under about 10 tasks between single runs are unresolved.
How fast is it, and what hardware does it need?
The README reports a median of 115 ms (299 ms at the 95th percentile) per JevBench question on an RTX 3090 under WSL2, and a 153 ms warm median on an M3 Pro Mac for prompts under 300 tokens. It runs on CUDA, Apple silicon (MPS or MLX) or CPU.
How is it different from a normal LLM?
Its language-modelling head is removed. A pointer head of about a million parameters scores each option against the hidden state at the <answer> position in one forward pass, with no generation or decoding loop, and a softmax turns the scores into per-option probabilities.
Can it answer questions about images?
Yes, optionally. Qwen3.5-2B-Base is natively multimodal; started with --vision, the server keeps the vision tower and a request may carry base64 images as part of the state, using the same v19 checkpoint.
What license is it released under?
Apache-2.0, for both the GitHub repository and the Hugging Face weights.