Overview
Microsoft-Decision-1 is a decision-scoring model that Microsoft announced on October 9, 2026, in a post by Achint Srivastava of the Office of the CTO. It is built for routing, classification, prioritization, verification and workflow control. Given a fixed set of answer options — yes/no, multiple choice or a rating — it returns a calibrated probability for each option through a structured API call, and it can also grade AI responses and agent actions against a rubric. It does not generate free text.
The model is post-trained from Alibaba's open-weight Qwen3.5-9B, and Microsoft says it will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI models. The Microsoft Foundry catalog lists it as Version 1, direct from Azure, generally available, with a 32,768-token context window, text input and JSON output. Microsoft's post does not link open weights for the model itself.
In Microsoft's own benchmarking it had the lowest latency measured: P50 latency about 35x faster than GPT-6 Sol and 2.5x faster than the runner-up, H2O-Lightning-4B v1.1. Microsoft reports the highest accuracy in a 36-benchmark comparison of nearly 150,000 questions held out from training, against Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B and Strands Decider 2B. Across eight perturbation types it changed its decision on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased or options were reversed or shuffled.
Microsoft also reports internal results. On more than 10,000 XBOX Research feedback items its quality was competitive with GPT-6 Sol while running over 14x faster and 200x cheaper; for the Copilot team it was competitive with GPT5.6 Luna and 100x faster; and in Microsoft Discovery it was 46x more consistent than an LLM-based score at three times the speed. Input tokens cost $0.042 per million and output tokens are free.
| Released | 2026-10-09 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Context | 32,768 tokens |
| Architecture | Qwen3.5-9B post-trained by Microsoft for fast, single-pass decision scoring; returns per-option probabilities as JSON instead of generated text |
| Modalities | Text |
| Status | Generally available |
Pricing
| Input | $0.042 / 1M tokens |
|---|---|
| Output | $0 / 1M tokens |
Output tokens are free.
Strengths
- P50 latency about 35x faster than GPT-6 Sol and 2.5x faster than H2O-Lightning-4B v1.1 in Microsoft's benchmarking
- Highest accuracy in Microsoft's 36-benchmark comparison of nearly 150,000 held-out questions
- Stable answers: 1.3% average decision changes across eight perturbation types, and zero flips when options are paraphrased, reversed or shuffled
- Calibrated probability for every answer option, returned as JSON from one structured call
- $0.042 per million input tokens, with output tokens free
Best for
- Reach for it to route requests, tickets or tasks to the right queue, team or model inside an agent pipeline.
- Reach for it for agent guardrails and content-safety screening, acting on the option probabilities it returns.
- Reach for it to grade AI responses or agent actions against a rubric at a fraction of an LLM judge's cost.
- Look elsewhere for open-ended generation, conversation, translation or summarization, and do not use it as the sole decision-maker in consequential decisions about people — the Foundry catalog rules out both.
How to access
| Provider | Model ID |
|---|---|
| Microsoft Foundry ↗ | Microsoft-Decision-1 |
FAQ
What is Microsoft-Decision-1 used for?
Microsoft-Decision-1 scores a fixed set of answer options and returns a calibrated probability for each. Microsoft built it for routing, classification, prioritization, verification and workflow control, and the Foundry catalog lists AI-output evaluation, agent guardrails, content-safety screening and document relevance as uses.
How much does Microsoft-Decision-1 cost?
Input tokens cost $0.042 per million, and output tokens are free, according to Microsoft's launch post.
Is Microsoft-Decision-1 open-weight?
No weights are linked. Microsoft-Decision-1 is post-trained from the open-weight Qwen3.5-9B, but Microsoft offers it as a hosted model in Microsoft Foundry and its launch post gives no Hugging Face or GitHub page for the model.
How fast is Microsoft-Decision-1?
In Microsoft's benchmarking its P50 latency was about 35x faster than GPT-6 Sol and 2.5x faster than the runner-up, H2O-Lightning-4B v1.1. These are the authors' own measurements.