Jev AI

Model guide · hosted decision model · published October 11, 2026

Microsoft-Decision-1 vs Jev

Microsoft’s first decision model answers the same typed questions as Jev. Microsoft says it is the most accurate and the fastest. We ran both on the same day, the same route and the same 740 decisions.

Microsoft-Decision-1 is a Qwen3.5-9B model post-trained to return a probability for every option of a noul, choice or score question, without generating text. It is hosted in Microsoft Foundry and on OpenRouter; the weights are not public. Microsoft’s launch charts put it first on accuracy across 36 benchmarks and first on speed, with Jev a point behind on accuracy and ahead on calibration. Our test agrees on direction and on size: a small lead for Microsoft on both of our tasks.

92.7% vs 91.0%our 300 HaluEval verdicts, Microsoft-Decision-1 against Jev, same day and route
95.0% vs 93.6%our 440 BFCL tool-routing decisions, same pair
83.5% vs 82.3%Microsoft’s 36-benchmark average; Jev leads its calibration score, 93.7 to 92.2
Text onlyan image sent as an image part was read as 6,200 text tokens

Microsoft-Decision-1 at a glance

DeveloperMicrosoft
Released9–10 October 2026, in Microsoft Foundry and on OpenRouter; OpenRouter serves the snapshot microsoft-decision-1-20261009
Base modelQwen3.5-9B, post-trained for single-pass decision scoring; Microsoft says it will rebase on other models, including MAI and OpenAI models
WeightsNot released. Hosted only, in Foundry and through OpenRouter
InputsText or JSON state. Images are not read: sent as an image part, the picture was tokenized as text in our test
Question typesnoul, choice and score, with a probability for every option; no generated text or rationale
Foundry endpoint<resource>/providers/microsoft/v1/systemone; the request’s model is your deployment name, not the model name
Deployment typesGlobalStandard, and DataZoneStandard in selected regions to keep inference in one data zone
AuthenticationMicrosoft Entra ID (recommended) or an API key

What Microsoft reports

36 benchmarks kept blind from training, 147,137 questions. Microsoft updated its post after launch to add Jev. These are the vendor’s numbers, not ours.

ModelAverage accuracyMedian latencyCalibration (100 = perfect)
Microsoft-Decision-183.5%85 ms92.2
Jev 1.13.082.3%240 ms93.7
Quyet-1.0-Large81.9%380 ms93.1
Surogate Rune 26B-A4B79.7%380 ms91.8
GPT-6 Luna Decisions79.4%300 ms89.9
deck-31B77.8%400 ms83.5
H2O-Lightning-4B v1.177.2%210 ms91.8
Strands-Decider 2Banswered only 23 of the 36 benchmarks54.8%——

Microsoft-Decision-1 p95 latency 125 ms; GPT-6 Sol, a generative reference, 3,010 ms median. Across eight kinds of rewording, reordering and formatting noise, decisions flipped on 1.3% of perturbations, and never when options were paraphrased, reversed or shuffled. Tested on 5,250 harmful, jailbreak and prompt-injection requests across 11 benchmarks. Latency for every model except Microsoft-Decision-1 comes from JevBench v1.6.1; Microsoft measured its own model through Foundry, so the speed gap mixes a model difference with a serving difference. Internal uses Microsoft names: Xbox Research labelling 10,000+ player comments, Copilot response-quality grading, incident-response retrieval and Microsoft Discovery experiment grading.

Our test: same day, same route, same items

Microsoft-Decision-1, Jev and Cloudflare’s new Clef-Omni, each through OpenRouter Decisions API from a laptop in Asia, 4 requests at a time, 2026-10-11.

HaluEval judging accuracy

  • Microsoft-Decision-192.7%
  • Jev91.0%
  • Clef-Omni76.0%

Rejects hallucinated answers

  • Microsoft-Decision-191.3%
  • Jev87.3%
  • Clef-Omni57.3%

BFCL tool routing accuracy

  • Microsoft-Decision-195.0%
  • Jev93.6%
  • Clef-Omni86.1%

Declines when no tool fits

  • Microsoft-Decision-192.5%
  • Jev88.8%
  • Clef-Omni74.6%
ModelJudge accuracyAccepts · rejectsJudge ECERouting accuracyRight tool · declinesRouting ECEMedian per call
Microsoft-Decision-192.7%94.0% · 91.3%0.05595.0%98.0% · 92.5%0.028688 ms
Jev (jev-1.13.0)91.0%94.7% · 87.3%0.05793.6%99.5% · 88.8%0.046623 ms
Clef-Omni76.0%94.7% · 57.3%0.15786.1%100.0% · 74.6%0.086546 ms

What we found

  • Microsoft-Decision-1 led Jev on both tasks: by 1.7 points on judging and 1.4 points on routing. That matches the size of Microsoft’s own lead, about a point, and is within what a second run could move.
  • The lead comes from saying no. It rejected 91.3% of hallucinated answers to Jev’s 87.3%, and declined 92.5% of requests with no fitting tool to Jev’s 88.8%. When a tool did fit, Jev picked it slightly more often, 99.5% to 98.0%.
  • Calibration was close: expected calibration error 0.055 against 0.057 on judging, 0.028 against 0.046 on routing. Microsoft’s own charts give Jev the edge here; on our routing set Microsoft’s model was slightly better.
  • Through OpenRouter from Asia, Jev was slightly quicker: 623 ms against 688 ms median per verdict. The relay and distance dominate both; Microsoft’s 85 ms is a Foundry figure measured close to the model.

How we tested

  • Judging: 150 HaluEval QA items, a correct and a hallucinated answer each, 300 Noul verdicts; the wording from our Jev-as-a-Judge test.
  • Routing: 200 BFCL requests with one right tool and 240 where none fits, one Choice per request with a “none” option; as in our Jev agent test.
  • Route: every model through OpenRouter’s Decisions API on 2026-10-11, four requests at a time, so latency and availability were the same for all. Microsoft-Decision-1 resolved to microsoft/microsoft-decision-1, Jev to jev-1.13.0.
  • Limits: two public datasets, one wording written for Jev, one run each. We did not rerun Microsoft’s 36 benchmarks.

Call it

Foundry for production on Azure, OpenRouter to try it next to other decision models.

curl "$AZURE_ENDPOINT/providers/microsoft/v1/systemone" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $AZURE_ENTRA_TOKEN" \
  -d '{
    "model": "'"$DEPLOYMENT_NAME"'",
    "state": "The API returns 500 on every call.",
    "questions": {
      "team": { "type": "choice", "instructions": "Which team should handle this ticket?",
                "criteria": { "billing": "Charges, invoices, and refunds",
                              "engineering": "Bugs, errors, and outages",
                              "support": "How-to and account questions" } },
      "urgent": { "type": "noul", "instructions": "Does this need a response today?" }
    }
  }'
curl https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "microsoft/microsoft-decision-1",
        "state": "The API returns 500 on every call.",
        "questions": { "urgent": { "type": "noul", "instructions": "Does this need a response today?" } } }'

In Foundry, the model field is your deployment name; the response’s model names the underlying model. Choose DataZoneStandard to keep inference in one data zone, or GlobalStandard for global capacity at the price of more variable latency. The request shape is the one Jev uses, so moving a workload is mostly an endpoint and credential change; see where to run Jev for the equivalent Jev routes.

Microsoft-Decision-1 or Jev?

Pick Microsoft-Decision-1 if

  • You are on Azure and want the model inside your Foundry project, with Entra ID and a data-zone deployment.
  • Your decisions are mostly gates and checks where saying no matters, the side where it led in our test.
  • You want the lowest latency close to the model, which Microsoft reports.

Pick Jev if

  • You want the calibration Microsoft’s own charts rank first.
  • You already use TypeSafe’s SDKs or Jev AI’s API and do not want to move a working integration for a point or two.
  • You want to try it now with no Azure account: the playground runs Jev in the browser.

Related

Clef-Omni for images and audio · pplx-decider · Jeeves · local models on Ollama · JevBench

Microsoft-Decision-1 FAQ

What is Microsoft-Decision-1?

Microsoft’s decision model, released on 9 and 10 October 2026 in Microsoft Foundry and on OpenRouter. It is Qwen3.5-9B post-trained to answer typed questions in one pass: noul (yes or no), choice and score, each with a probability for every option. It does not write text. Microsoft positions it for routing, classification, grading, agent controls and content filtering.

Is Microsoft-Decision-1 better than Jev?

On Microsoft’s own 36-benchmark comparison it averages 83.5% against Jev’s 82.3%, while Jev scores higher on calibration, 93.7 against 92.2. In our test the same day through the same route, it scored 92.7% against 91.0% on 300 HaluEval verdicts and 95.0% against 93.6% on 440 tool-routing decisions. Both gaps are small: a point or two on one run each. Benchmark Heaven’s independent JevBench, which measured it natively through Foundry on 1,500 decisions, ranks it #6 of 27 on its API board at 69.1 against Jev’s #4 at 71.5, with calibration 84.4 against 90.6.

Is Microsoft-Decision-1 on Hugging Face?

No. Microsoft has not released the weights; the model runs only as a hosted service in Microsoft Foundry and through OpenRouter. For open decision models you can run yourself, see Jeeves, Kev, Clef-Omni and the Ollama models.

How fast is Microsoft-Decision-1?

Microsoft reports an 85 ms median and 125 ms p95 through Foundry, against 240 ms for Jev in JevBench v1.6.1; JevBench’s own native measurement of Microsoft-Decision-1 found a 459 ms median on its speed subset. Through OpenRouter from Asia, which adds a relay and long-distance hops, we measured 688 ms median per verdict against 623 ms for Jev on the same route; at that distance the network, not the model, sets the pace.

Can Microsoft-Decision-1 read images?

No. It takes text or JSON state. When we sent a photo as an image part through OpenRouter, the request was accepted but the image was counted as about 6,200 text tokens and the answer was wrong. For images, audio or video, use a multimodal decision model such as Clef-Omni.

How do I call it?

In Foundry, deploy Microsoft-Decision-1 from the model catalog and POST to <resource>/providers/microsoft/v1/systemone with your deployment name as the model, authenticated with Entra ID or an API key. Through OpenRouter, POST to the Decisions API with model microsoft/microsoft-decision-1. The request body is the same state and questions shape Jev uses.

Does my Jev code work with Microsoft-Decision-1?

The request and response shapes match: state, a questions map, and answers keyed by your question ids with probabilities. Change the endpoint, the credential and the model name. Re-check thresholds, because the two models are calibrated differently.

About this page

Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and publishes measurements of it and its alternatives. We are not affiliated with Microsoft or TypeSafe.

How. Facts are from Microsoft’s announcement and Foundry documentation; the vendor table is the data behind Microsoft’s published charts. Our numbers come from the scripts and raw result files in our repository, run on 2026-10-11.

When. Published October 11, 2026. Microsoft says it will update the model and rebase it on other families; we will rerun when the served snapshot changes.

Sources

Microsoft and Microsoft-Decision-1 are Microsoft’s; Jev is TypeSafe AI’s. Neither is affiliated with Jev AI.