Microsoft Ships Decision-1, a Qwen 9B Decision Model at Jev's Exact $0.042 Price
TL;DR
Microsoft put a decision model into Foundry and OpenRouter on October 9. Microsoft-Decision-1 is Qwen3.5-9B post-trained to read a state plus a set of typed questions and return a calibrated probability for every option, in one pass, with no text. It uses the same three primitives as TypeSafe's Jev (noul, choice, score), the same systemone endpoint name, and the same price: $0.042 per million input tokens, output free. On Microsoft's own 36-benchmark panel it scores 83.5% to Jev's 82.3%, posts an 85 ms median latency, and lands third on calibration. Every number is Microsoft's, the latency setups differ, and there are no weights.
What shipped
The model is live in the Microsoft Foundry catalog and on OpenRouter as microsoft/microsoft-decision-1, where the single endpoint is Azure-backed, the context window is 32,768 tokens, and the modality field reads "text->decisions". Input costs $0.042 per million tokens. Output is free, because nothing is generated to bill.
Under the hood it is a post-trained Qwen3.5-9B, which Microsoft says it will "soon rebase" on its own MAI models and on OpenAI models. The announcement is signed by Achint Srivastava, VP of Software Engineering in the Office of the CTO, and the Microsoft Learn docs carry an October 7 date, two days ahead of the post.
If you have never used one of these: an LLM judge is an essay exam, writing its reasoning one token at a time while you parse the verdict out of the prose. A decision model is a bouncer with a clipboard. It looks at you once, checks the list, and hands back a number, and that single forward pass is where both the speed and the zero output bill come from.
The API is Jev's API
Microsoft's docs describe the request pattern as "similar to a system-one API", which is TypeSafe's term for Jev. The resemblance goes further than the concept. Jev's endpoint is https://api.typesafe.ai/v1/systemone; Microsoft's is <your-foundry-resource>/providers/microsoft/v1/systemone. Both take a state and a questions map, each question has a type, instructions, and optional criteria, and both return an answers object keyed by your question names.
The three question types match TypeSafe's noul, choice, and score primitives by name. A noul question returns a probability from 0 to 1. A choice returns one selected option plus a probability for every option. A score places the item on an ordered scale you define, and the returned value is the probability-weighted average of the level indexes, so a four-level scale can hand you 2.4.
{
"model": "<deployment-name>",
"state": "The API returns 500 on every call.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, invoices, and refunds",
"engineering": "Bugs, errors, and outages",
"support": "How-to and account questions"
}
}
}
}
You can stack several questions against one state in a single request, say a choice for routing, a score for severity, and a noul for repeat contact. The docs recommend an abstention option such as "cannot tell", randomized option order in testing, and a confidence threshold you set from the cost of your own false positives. Deployments come in DataZoneStandard and GlobalStandard flavors, and the request's model field is your deployment name, not the model name.
Microsoft's scoreboard
Microsoft built a panel of 36 benchmarks, 147,137 questions in total, mixing public and private sets it says were kept blind from training and spanning routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning, and safety. It then took several of the top entries on the JevBench leaderboard (positions as of October 8) and ran them through the same panel. The chart data embedded in the post gives the averages.
The post carries an editor's note: it was updated after publication "to add benchmarks for Jev on accuracy and calibration". The original chart managed to leave out the model the whole category is named after. Once added, Jev took second on accuracy, 1.2 points back, and first on calibration.
Latency
Microsoft-Decision-1 reports an 85 ms median and a 125 ms p95, which the chart footnote says was "measured through Foundry in the same region". The other models' figures are JevBench v1.6.1 adjusted medians, checked October 7, which is a different harness and a different network path. That is where the headline "2.5 times quicker than H2O-Lightning-4B v1.1, the runner-up, and 35 times quicker than GPT-6 Sol" comes from.
Calibration and robustness
Calibration is where Microsoft does not win. Its chart scores Jev 1.13.0 at 93.7, Quyet-1.0-Large at 93.1, and Microsoft-Decision-1 at 92.2, with Surogate Rune and H2O-Lightning-4B at 91.8, GPT-6 Luna Decisions at 89.9, and deck-31B at 83.5. A score of 100 would mean the model's confidence is a perfect odds-maker: of everything it marks 70%, exactly seven in ten are right. Microsoft's docs add that calibration "is strongest on familiar task types", which is a polite way of saying your own labeled set is the real test.
On robustness, Microsoft perturbed the same request eight ways (paraphrased state, changed option descriptions, reordered choices, renamed keys, formatting noise) and counted decision flips: 1.3% on average, with zero flips when option descriptions were paraphrased or options reversed or shuffled. Safety testing covered 5,250 requests across 11 benchmarks of harmful content, jailbreaks, and prompt injection; the post reports it "refused harmful behavior while retaining a high degree of utility" and gives no rate.
The Xbox cost run
The most concrete internal test comes from Xbox Research, which pushed more than 10,000 open-ended pieces of feedback from surveys, Steam, and X through the model to sort them into researcher-defined themes. Microsoft calls the quality "competitive" with GPT-6 Sol at over 14x the speed and 1/200th the cost, and the chart data backs the ratio: a racing title's 753 reviews took 2,800 ms per review on Sol and 159 ms on Decision-1 ($1.83 against $0.009 for the run), a shooter's 341 reviews 2,600 ms against 188 ms ($1.03 against $0.005), and 1,030 livestream posts 2,600 ms against 143 ms ($2.92 against $0.013).
Scaled to a million texts, Microsoft estimates about $2,434 on Sol and about $11 on Decision-1. The footnote marks the costs as estimated and explains the gap: both models read the identical brief, so input tokens match, and only Sol pays for output. The Copilot team reports the model "competitive with GPT5.6 Luna and 100 times faster" at grading chat and agent responses, and Microsoft Discovery's replanning loop scored it 46 times more consistent than an LLM-based rubric score at three times the speed.
A category that is 25 days old
Jev came out of stealth on September 16. Since then the shelf has filled fast: Cloudflare's Clef and Clef-flash (Apache 2.0, Qwen3.8-27B and Qwen3.5-9B backbones, with Clef-flash now listed at $0.038 per million input tokens on Workers AI), the one-home-GPU Jeff models, H2O-Lightning-4B (Apache 2.0, built on Qwen3.5-4B, uploaded October 2), AWS's Strands-Decider, OpenAI's GPT-6 Luna Decisions mode, and Quyet-1.0-Large, a Hugging Face upload from user chinhnc that was sitting at #1 on JevBench on October 8. Most of the open ones are small Qwen fine-tunes, a point one Hacker News commenter made and Simon Willison summarized as "I guess the fear of Chinese models is finally subsiding."
Microsoft's post went up the same day TypeSafe announced an $870 million Series A at a $7.5 billion valuation, led by Andreessen Horowitz, with the claim that a third of the Fortune 500 is already using Jev. Nothing says "we noticed" quite like shipping the look-alike on funding day. The practical difference for you is that Microsoft's model is hosted only: no weights, no license, no Hugging Face page, which puts it in Jev's bucket rather than Clef's.
Caveats
- Every number is Microsoft's. The 36 benchmarks are a Microsoft-chosen mix of public and private sets, the comparison models were picked from JevBench, and nobody outside Microsoft has reproduced the panel. Treat 83.5% as a claim, not a result.
- The latency comparison is cross-harness. Decision-1 was timed in-region through Foundry; the others are JevBench adjusted medians from a different setup. Microsoft also measured its own model with a p95; the others get p50 only.
- The margins are thin. 83.5 versus 82.3 versus 81.9 is 1.6 points across three models on a private panel, and Jev was added to the charts after publication. On calibration Microsoft is third.
- The model will change under the name. Microsoft says it will rebase on MAI and OpenAI models "soon". Pin a deployment and re-run your labeled set when the version moves.
- Smaller window than Jev. OpenRouter lists 32,768 tokens of context; TypeSafe's docs give Jev 64k per request, with 32k for the state. The docs also warn of wording sensitivity, no rationale, and misses on subtle harmful content when used as a filter.
Key Takeaways
- Microsoft-Decision-1 is a Qwen3.5-9B post-trained for single-pass decision scoring, live in Microsoft Foundry and on OpenRouter since October 9, at $0.042 per million input tokens with free output.
- The API mirrors Jev's: a
systemoneendpoint, a state plus named questions of typenoul,choice, orscore, and probabilities back per option. Porting between the two looks like little more than a URL and auth change, though the response fields are not documented as identical. - On Microsoft's 36-benchmark, 147,137-question panel it averages 83.5% against Jev's 82.3% and Quyet-1.0-Large's 81.9%, with an 85 ms median latency; on calibration it ranks third at 92.2 behind Jev's 93.7.
- Xbox Research's labeling run put it at 143 to 188 ms and about $11 per million texts, against 2.6 to 2.8 seconds and about $2,434 on GPT-6 Sol, in Microsoft's own estimates.
- It is hosted only, with no weights or license, and Microsoft plans to swap the base model. Validate on your own labeled data and pin the deployment.
Sources: Microsoft Command Line announcement, Microsoft Learn docs, Foundry model catalog, OpenRouter listing, JevBench leaderboard, TypeSafe Jev docs, TypeSafe models and pricing, TypeSafe Series A post, Hacker News thread, TestingCatalog, h2o-lightning-4b model card, OpenAI GPT-6 Sol pricing