█

AI/TLDR

Strands Decider 2B

A 1.9B Apache-2.0 decision model from AWS Strands Labs that picks between options or rates on a scale with a calibrated confidence, in a single forward pass on local hardware, released October 2026.

Strands Decider (open decision models)Open weights
Released
1 Oct 2026
Parameters
1.9B (Qwen3.5-2B-Base torso with a rank-16 LoRA adapter and a pointer head of about 1M parameters)
License
Apache-2.0
Coverage
1 story

Overview

Strands Decider 2B is an open decision model, or "system one" model, published by AWS's Strands Labs on October 1, 2026, with the code on GitHub under Apache-2.0 and the weights on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19. Unlike an LLM it does not generate text: it picks between a set of options or rates something on a scale, and every answer carries a confidence. It is built for the decisions inside agentic workflows made with the Strands Agents SDK, such as model routing, tool selection, argument checking, triage, guardrails and evaluations.

Strands Decider 2B architecture diagram: on the left a decoder LLM whose LM head and generated text are marked discarded; on the right the same Qwen3.5-2B-Base torso with a rank-16 LoRA reads state, question and numbered options, and a pointer head compares the hidden state of each option with the hidden state at the answer token to produce logits, then a softmax gives the answer and per-option probabilities.
Keep the pretrained torso, discard the language-modelling head, score each option with a pointer head.AWS Strands Labs ↗

The design keeps the torso of Qwen3.5-2B-Base, discards its language-modelling head and adds a pointer head of about a million parameters. The head scores each option by comparing the hidden state at the <answer> position with the hidden state at that option's last token, and a masked softmax turns the scores into per-option probabilities. The torso is adapted with a rank-16 LoRA adapter. Because the head has no per-option parameters, label sets come from the request rather than the weights and there is no cap on how many options a question may carry. Three question types share the same head: noul (yes/no), choice (one of N) and score (an ordered rubric).

Scatter chart of JevBench accuracy against Brier score for each Strands Decider 2B generation: the default recipe moves from v5 near 0.64 accuracy and 0.46 Brier score to v19 near 0.72 accuracy and 0.34 Brier score, with rejected experiments in orange and other systems such as frozen Qwen3.5-2B and Qwen3.5-4B in green.
Accuracy and calibration across the training trajectory on JevBench's public set.AWS Strands Labs ↗

The released checkpoint is v19. On JevBench's 231 public tasks it scores 0.723 (167 of 231), with a Brier score of 0.342 and an expected calibration error of 0.052, and 1.000 / 0.875 / 0.505 on the repository's easy, standard and hard tiers. The launch post places it 3rd of 33 in the 2B class on JevBench's public set. The README reports a median latency of 115 ms (299 ms at the 95th percentile) per JevBench question on an RTX 3090, and a 153 ms warm median on an M3 Pro Mac for prompts under 300 tokens. On short classification tasks it has never seen, answers at a confidence of 0.9 or more are right about 95% of the time.

It installs with pip install strands-decider and runs from a CLI or as a local server with a /v1/systemone endpoint, on CUDA, Apple silicon (MPS or MLX) or CPU. With --vision the server keeps Qwen3.5-2B-Base's vision tower so a request can carry images. The repository also holds the training recipe (about 11 hours on one RTX 3090, or 1 hour 10 minutes on eight H100s), the data sources and a preregistered record of every training run.

Released2026-10-01
LicenseApache-2.0
WeightsOpen weights
Parameters1.9B (Qwen3.5-2B-Base torso with a rank-16 LoRA adapter and a pointer head of about 1M parameters)
ArchitectureQwen3.5-2B-Base decoder torso with its language-modelling head discarded and replaced by a pointer head that scores each option against the hidden state at the <answer> position; one forward pass, no generation
ModalitiesText, Vision

Benchmarks

Strands Decider 2B (v19) against the nearest open systems on JevBench's 231 public tasks, as published in the repository's evaluation notes. v19 is Strands Labs' own harness run; the others come from the benchmark's per-task data. Tiers are the repository's split of the public tasks.

BenchmarkStrands Decider 2B (v19)decider-2bOpen-Jev 2Bjeff 400MLaya 421MOpen-Jev 9B
Overall (231 public tasks)0.7230.710.6450.6280.5840.775
Easy tier11110.9581
Standard tier0.8750.8470.7640.750.6940.903
Hard tier0.5050.4950.4140.3870.3510.595
routing_hard110.80.60.21
intent0.91710.7080.7080.8331
adversarial0.6670.6670.510.6671
long_policy0.3680.3160.2630.1050.2110.474
multi_hop0.4440.5560.50.3890.3330.667
policy0.750.8330.9170.9170.751
trap0.87510.8750.501
judge_hard0.5880.5290.5290.4710.4120.765
temporal_numeric0.2670.3330.0670.40.3330.133

Comparison source ↗

This model's scores

  1. JevBench v1, public set (231 tasks)72.3%
  2. JevBench, easy tier100%
  3. JevBench, standard tier87.5%
  4. JevBench, hard tier50.5%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Strengths

  • 0.723 on JevBench's 231 public tasks (167 of 231), with a Brier score of 0.342 and an expected calibration error of 0.052
  • Every task in JevBench's easy tier answered correctly (1.000), and 0.875 on the standard tier
  • Median 115 ms per JevBench question on an RTX 3090, and it also serves on an Apple-silicon Mac or on CPU
  • Many questions about one text are cheap: the text is read once and each question adds only its own tokens
  • Calibrated confidences: on unseen short classification tasks, answers at 0.9 confidence or more are right about 95% of the time
  • Apache-2.0 code and weights, with the full training recipe retrainable in about 11 hours on one RTX 3090

Best for

  • Reach for it to make the rote decisions inside a Strands agent, such as which tool to call next or whether a tool call's arguments are valid, and leave the hard decisions to an LLM.
  • Reach for it to route an incoming request to the right team or queue, or to pick which LLM should handle a task, locally and in about a tenth of a second.
  • Reach for it for guardrails and evaluations: grounding checks, policy classification and scoring model outputs at low cost, acting only on high-confidence answers.
  • Look elsewhere for the hard tier of JevBench, where it scores 0.505 and the 9B Open-Jev scores 0.595, or for anything that needs generated text.

How to access

ProviderModel ID
Hugging Face (weights) ↗StrandsAgents/strands-decider-2B-hobson-v19
PyPI (strands-decider CLI and server) ↗—

FAQ

What is Strands Decider 2B?

An Apache-2.0 decision model from AWS Strands Labs, released October 1, 2026. Instead of generating text, it picks between options (yes/no or one of N) or rates something on a scale, and returns a confidence with every decision. It is meant for the decisions inside agentic workflows built with the Strands Agents SDK.

How accurate is it?

The released v19 checkpoint scores 0.723 (167 of 231) on JevBench's public set, with a Brier score of 0.342 and an expected calibration error of 0.052. The launch post places it 3rd of 33 in the 2B class. The README notes that six retrains of an earlier recipe varied by 3.2 tasks, so differences under about 10 tasks between single runs are unresolved.

How fast is it, and what hardware does it need?

The README reports a median of 115 ms (299 ms at the 95th percentile) per JevBench question on an RTX 3090 under WSL2, and a 153 ms warm median on an M3 Pro Mac for prompts under 300 tokens. It runs on CUDA, Apple silicon (MPS or MLX) or CPU.

How is it different from a normal LLM?

Its language-modelling head is removed. A pointer head of about a million parameters scores each option against the hidden state at the <answer> position in one forward pass, with no generation or decoding loop, and a softmax turns the scores into per-option probabilities.

Can it answer questions about images?

Yes, optionally. Qwen3.5-2B-Base is natively multimodal; started with --vision, the server keeps the vision tower and a request may carry base64 images as part of the state, using the same v19 checkpoint.

What license is it released under?

Apache-2.0, for both the GitHub repository and the Hugging Face weights.