Model Releases 2026-10-11

Microsoft-Decision-1 Is a 9B Model That Only Picks From a List -- for $0.042 per Million Input Tokens, With Output Free

Microsoft released Microsoft-Decision-1 on October 9: a model post-trained from Qwen3.5-9B that returns a calibrated probability for each of a fixed set of options instead of generating text. It is aimed at the many small yes/no, routing and grading steps inside agent workflows, and Microsoft's own teams describe using it for feedback triage, response grading and incident retrieval.

The announcement, by Achint Srivastava, VP of Software Engineering in Microsoft's Office of the CTO, went up on October 9, 2026. The model is available in Microsoft Foundry and on OpenRouter.

What a decision model does

"Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on," Srivastava writes. Given a fixed set of answer options, Microsoft-Decision-1 returns "a calibrated probability score for each option," covering yes/no, multiple-choice and rating questions as well as "rubric-based grading of AI responses and agent actions." Microsoft says it "post trained Qwen3.5-9B" and "will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI."

Why latency is the selling point

The post frames latency as a compounding cost: "adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow." Microsoft reports the highest accuracy in its "36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training," and measured it as "2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol." All of these are Microsoft's own benchmarks.

Two design points matter for production. Consistency: requests were perturbed eight ways, and the model "changes its decision on 1.3% of perturbations on average," with zero flips when options are reordered. Calibration: "a 90% prediction should be right about nine times out of 10," because applications use the probability "to decide when to act, defer, or ask for review."

How Microsoft teams are using it

  • Xbox Research sorted more than 10,000 pieces of open-ended feedback from surveys, Steam and X into researcher-defined themes, finding it "competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive."
  • Copilot's quality team, grading chat and agent responses, found it competitive with GPT5.6 Luna and "100 times faster."
  • On-call engineers use it for knowledge retrieval across logs, tickets, calls and messages during live incidents.
  • Microsoft Discovery uses it to grade experiments in an adaptive replanning loop, where it scored "46 times more consistent than the LLM-based score."

Pricing. "Input tokens cost $0.042 USD per million tokens. Output tokens are free" – free because the output is a handful of probabilities, not text.

What remains uncertain. There is no independent benchmark yet. The model also inherits Qwen3.5's base, so procurement teams with country-of-origin rules will want to know when the promised rebase lands.

Analysis: this is the same routing argument made in this week's DeepSeek debate, packaged as a product. Most steps in an agent loop are not open-ended reasoning but choices: continue or stop, which tool, accept or revise. Moving those to a cheap, calibrated classifier cuts latency and cost, and a probability gives a natural threshold for human review. Developers should test it on their own decision points, since calibration on Microsoft's benchmarks does not guarantee calibration on theirs.

Microsoft-Decision-1, post-trained from Qwen3.5-9B, returns calibrated probabilities over fixed options for routing, grading and agent control at $0.042 per million input tokens with free output; Microsoft's own benchmarks show it 35 times faster than GPT-6 Sol, and internal teams use it for feedback triage, response grading and replanning.