Microsoft's latest model isn't built to write or chat. It's built to pick, fast, and Redmond says it decides far quicker than the competition.
Microsoft-Decision-1 Is a Model That Chooses, Not Chats
Microsoft introduced Microsoft-Decision-1 on October 9, and it breaks from the pattern of recent model launches. It doesn't write essays or hold conversations. Given a situation and a fixed set of options, it returns a probability for each one instead of generated text. Satya Nadella announced it on X. It's live in Microsoft Foundry now, with OpenRouter coming soon.
That narrow focus is the point. Most models try to do everything. This one does a single job: score a set of choices and hand back an answer software can act on right away. For developers wiring up routing, classification, or agent controls, that's a different kind of tool than a chatbot.
How Microsoft-Decision-1 Scoring Works
Decision-1 accepts yes/no questions, multiple-choice options, and rating scales. It also handles rubric-based grading, where it scores an AI response or an agent action against a set of criteria. Then it returns calibrated probabilities, which means the numbers are meant to reflect real odds rather than raw confidence scores.
The model is a small one. Microsoft post-trained Qwen3.5-9B for single-pass scoring, and says it plans to rebase the system on other models later, including Microsoft AI and OpenAI technology. A compact model matters here because decision scoring often runs at the center of a loop. If an agent has to choose a path every few seconds, the cost and speed of each call add up fast.
Microsoft points to concrete uses: routing requests, tagging data, prioritizing tasks, verifying outputs, and controlling agents. Those are unglamorous jobs. They're also the ones that make an AI product work or fall apart at scale.
What Microsoft Says About Decision-1's Speed
The speed claim is the headline. According to Microsoft, Decision-1 was 35 times faster than GPT-6 Sol at P50 latency in its own tests, and 4.5 times faster than the runner-up, Quyet-1.0-Large. The company also says the model posted the highest accuracy across a 36-benchmark evaluation covering nearly 150,000 questions held out from training.
Those numbers come from Microsoft, not an outside lab. The company's speed and accuracy claims haven't been independently verified. So treat the 35x figure as a vendor statement, not a settled fact. Speed matters more in some workloads than others. A benchmark run by the seller is a starting point, not a verdict. Is the speed real? Microsoft's own testing says yes, with a consistency figure to back it up. Across eight variations of the same request, the model's decision changed 1.3% of the time on average. It didn't flip when option descriptions were paraphrased, or when the options were reversed or shuffled. For anyone building a system that has to give the same answer twice, that consistency might matter more than how fast it gets there.
Inside Microsoft's Decision-1 Testing
Microsoft says it tested Decision-1 across its own products before shipping. According to the company, Xbox Research used it to label more than 10,000 feedback items. Internal trials also put it to work on incident routing and other workflows where a fast, repeatable choice beats a thoughtful paragraph.
The safety testing covered 5,250 requests across 11 benchmarks, aimed at harmful content, jailbreaks, and prompt injection. For a decision model, those risks look different than for a chatbot. There's no prose to abuse. But a model that controls an agent's next step can still be pushed toward a bad choice if the input is crafted carefully. Microsoft says the testing set out to check that the scorer holds up under those conditions.
What Microsoft-Decision-1 Means for Developers
If you build agents or data pipelines, Decision-1 is worth a look, but keep the claims in proportion. The appeal is a small, fast scorer that returns numbers instead of text, which is exactly what routing and verification steps need. The caution is that the strongest evidence so far comes from Microsoft's own benchmark runs rather than from an independent lab that had no stake in how the numbers turned out.
It's available now in Microsoft Foundry, with OpenRouter due soon, so you can test it against your own workload rather than trusting a benchmark page. Run your real inputs through it, check that the answers stay stable when the phrasing changes, and compare the cost per call to the LLM you'd otherwise use for the same job. A model built around making choices stands or falls on whether it picks the way you would, and the only comparison that settles that question is the one you run on your own data.






