Microsoft-Decision-1 Puts Agent Control to the Test
Microsoft-Decision-1 scores fixed choices for agent workflows. Its benchmark claims promise speed, but a latency dispute shows what developers need to test.
Written by AI. Dev Kapoor

Microsoft released Microsoft-Decision-1 on October 9 for a job many AI applications already do: choosing what happens next. The model scores defined options instead of generating an explanation. Microsoft offers it through Foundry and lists routing, verification and workflow control among its intended uses; its announcement says OpenRouter access is coming soon.
One proposed use is an agent checking its next step: continue, stop, retry or hand work to a tool, another model or a person. That is an attractive place to save time. A general-purpose model asked to judge every step may produce paragraphs when the application needs a choice. Decision-1 promises an answer that software can consume directly. The question for a developer is how often that answer should be allowed to move anything.
The Decision Was Here Before the Model
Classifiers have long turned inputs into labels or probabilities. What changed with language-model judging was the ability to describe criteria in ordinary language rather than first training a classifier on a substantial set of labeled examples. An O’Reilly essay on Jev, originally published on LinkedIn, describes that trade-off: an LLM judge can start from a prompt, while a conventional trained classifier needs examples and upkeep. The prompt route can also spend tokens generating reasoning when the application only needs a score.
TypeSafe AI’s Jev made the narrower job into a hosted interface before Microsoft’s launch. Its Choice, Score and Noul question types respectively select from options, rate ordered levels and estimate whether a statement is true. Jev returns typed answers and probabilities. Microsoft-Decision-1 follows that product pattern with fixed-option scoring built on Qwen3.5-9B, post-trained by Microsoft. Both are offered as hosted services, rather than as downloadable weights in the cited product descriptions.
The comparison helps locate Microsoft’s contribution. A team with labeled ticket history can still train a classifier for a stable queueing problem. A team changing its rubric frequently might prefer to state the options and criteria in a decision API. Microsoft packages that approach inside Foundry and explicitly pitches agent controls. Jev already shows that typed decisions are a product category, so the launch by itself does not settle which service makes better choices for any one team’s traffic. Meanwhile, other entrants include Cloudflare, Strands and H2O.ai. Deployment choices now include hosted APIs and models developers can run themselves; those arrangements carry different operational responsibilities as well as different invoices.
What the Scores Can Tell You
Microsoft says Decision-1 had the highest average accuracy in its comparison across 36 benchmarks and nearly 150,000 questions, with benchmarks kept blind from training. The average was 83.5%, versus 81.9% for the next model in the comparison. Those are Microsoft’s evaluation results, rather than an independent test on an adopter’s workload. They make a case for trying the model; an aggregate average cannot tell a support team how often an urgent ticket will be sent to the wrong queue.
Accuracy also answers a different question from whether probabilities support an action policy. If a model gives many wrong answers high scores, a threshold that automatically approves confident choices will let those errors through. If it assigns low scores to difficult cases, the application can route more of them to review, at the cost of more review work. An explicit “cannot tell” option can give missing or ambiguous cases somewhere to go. Developers still have to decide which cases qualify for automatic action, what happens on timeout and who receives an escalation.
That division of labor is visible in Vercel’s Jev implementation guidance. Its form-router example accepts a selected destination only when a configured confidence threshold is met and sends uncertain results or evaluation failures to a fallback model. The application keeps the mapping from an allowed destination to a receiving inbox. For an agent reviewing a proposed tool call, the equivalent design question is whether a score merely flags a call or actually permits execution. Permission checks and fixed business rules belong in application code, regardless of which model supplies the judgment.
A useful evaluation therefore starts with decisions people can check. Keep the question, permitted options, policy version and eventual correction alongside each case. Compare missed escalations with unnecessary escalations separately: they impose different costs on the people operating the workflow. Test ambiguous requests and cases outside the option list, then measure how many calls a chosen threshold sends to review. That procedure follows from the product’s design, not from a claim that one benchmark percentage can dictate a safe threshold everywhere.
There is a precedent for being wary of a single judge score. In a study of LLM marking in physics assessments, performance varied across structured questions, essays and scientific plots, and the researchers separated agreement on rankings from agreement on absolute marks. That study did not test Decision-1 or agent controls. It illustrates why an evaluation must match the question an application will actually ask: a model useful for ordering cases may still give scores unsuitable for an automatic approval cutoff.
The Latency Number Has a Serving Problem
Microsoft presents speed as another reason to put Decision-1 in the loop. Its comparison lists an 85-millisecond median for its own model and about 210 milliseconds for H2O-Lightning-4B. The methods behind those figures differ: Decision-1 was measured through Foundry, while the competitor figure uses a JevBench-adjusted median. A JevBench repository issue, opened October 11, says the board records a 29-millisecond raw median for H2O’s self-hosted model but displays roughly 210 milliseconds after an assumed adjustment of twice the measured time plus 150 milliseconds. The issue author asks for the estimate to be labeled or a hosted endpoint to be timed on the same basis as an API.
The 29-millisecond figure is not an end-to-end hosted-service substitute for Microsoft’s 85 milliseconds: it omits the assumed overhead the adjustment tries to represent. The displayed 210 milliseconds is likewise an estimate, not H2O’s raw measurement. For a developer deciding whether to put a call before every agent action, the relevant timing includes the serving setup, network and fallback path that application will use.
Microsoft has a coherent proposition: send a bounded question to a fast decision service, receive probabilities and let code branch on the result. Its own benchmark gives developers a reason to experiment. The operational test belongs at the branch. Record where Decision-1 sends real cases, where reviewers reverse it, how often it abstains or times out, and how long the whole handoff takes. The model can score “continue”; the application and its operators remain responsible for what continuing does.
More Like This
TypeSafe's Jev Bets on Faster Decisions for AI Agents
TypeSafe's Jev promises fast, cheap machine decisions. Its value depends on calibration, independent testing and whether existing tools already suffice.
Where Jev’s Cheap AI Decisions Work and Where They Fail
Two production benchmarks and a failed computer-use test show where Jev’s cheap structured decisions improve software, and where the costs can return.
Astra for Coding and the Treadmill of AI Agent Rebrands
A new AI coding agent launches, dev Twitter shrugs. Inside the launch-cycle loop and the benchmarks that would actually settle whether these tools work.
Meta Muse Glimmer 30B Runs in 14GB RAM via Unsloth
Meta's 30B coding agent fits in 14GB RAM thanks to Unsloth's dynamic 2-bit quantization. Here's what that buys you—and what it costs.
GPT-5.4's Schizophrenic Performance: A Model at War With Itself
ChatGPT 5.4 crushes quantitative tasks but fails basic reasoning. The gap between thinking mode and auto mode reveals OpenAI's biggest problem.
Scroll World: AI-Generated Scroll Animations Explained
Chase AI demos Scroll World, an open source skill that uses AI coding agents to build cinematic scroll-animated websites in a single prompt session.
Local AI's Inflection Point: Useful, Not Just Interesting
A panel of local AI builders at NVIDIA, Roboflow, Exo Labs, and r/LocalLLaMA maps where the movement stands—and what still needs solving.