Jev AI

Jev comparison · JevBench v1.5.4 · updated September 30, 2026

Jev vs Mercury Decide

Two hosted decision models that take the same request.

On this site

Jev 1.13.0

TypeSafe · System One model

The original System One model, with 64k of context and a public benchmark record.

72.1JevBench v1.5.4 rank #3 · Capability #1

vs

Alternative

Mercury Decide

Inception · Hosted diffusion decision model

Same schema, diffusion decoding, and up to 14 decisions per second by Inception’s count.

Not measured in JevBench v1.5.4

  1. JevSystem One: state plus typed questionsRequest schemaMercury DecideThe same System One schema
  2. JevReference answers (jev-1.13.0)Agreement with JevMercury Decide91.7% of 133 answers in our replay
  3. Jev64k tokensContext windowMercury Decide32k tokens
  4. JevMeasured in JevBench v1.5.4Independent benchmarkMercury DecideNot in JevBench v1.5.4 yet
  5. Jev0.65s median on the hosted APIThroughput claimMercury DecideUp to 14 decisions per second, per Inception
  6. JevGenerally availableAvailabilityMercury DecideEarly access through OpenRouter

Both models run on your Jev AI account with the same request: switch the model in the playground, or set model to mercury-decide in the API. Mercury Decide model guide →

JevBench v1.5.4, measured the same way

One benchmark measured all 112 systems under one method and ranked 106, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.

Composite score

#1 Cygnet73.7
#2 Winnow73.2
#3 Jev72.1
#4 JevK5 v0.371.9
#5 Plumb-4B71.6
#6 Jev-Omni71.5
#7 decider-4b71.3
#9 Imajev-4B70.4
#12 SemIf68.7
#15 Hopper67.5
#20 djev64.2
#93 Laya0.0

The top four ranked systems finish within 1.8 points of each other. The composite hides where systems actually differ, so the charts below break it apart. Benchmark Heaven’s headline Capability ranking, shown next to each score above, averages intelligence and calibration among Jev-class systems.

Mercury Decide was not measured in JevBench v1.5.4, so there is nothing to chart against Jev here. The analysis below uses the figures its authors publish.

One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.

Every system on the board

The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and Mercury Decide are highlighted.

ModelRankScoreMeasured latencyRuns on
Cygnet (blockbrain, frozen Gemma 4 12B)#173.70.04s raw, 0.23s adjustedSelf-hosted, RTX PRO 6000
Winnow-12B Q8 (EldanRing, Gemma 4 12B)#273.20.10s raw, 0.34s adjustedSelf-hosted, RTX 6000
Jev 1.13.0 (TypeSafe)#372.10.62s median / 0.67s p95Hosted API
JevK5 v0.3 (allebee, Qwen3.5-4B)#471.90.02s raw, 0.18s adjustedSelf-hosted, H100
Plumb-4B (crh225, JevK5 v0.2 + LoRA)#571.60.02s raw, 0.18s adjustedSelf-hosted, H100
Jev-Omni (akhilaaa3, Gemma 4 12B)#671.50.11s raw, 0.38s adjustedSelf-hosted, RTX 6000
decider-4b v2 (Mapika, Qwen3.5-4B)#771.30.03s raw, 0.20s adjustedSelf-hosted, RTX 5090
Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA)#970.40.04s raw, 0.23s adjustedSelf-hosted, RTX 5090
SemIf (Qwen3.5-4B)#1268.70.04s raw, 0.23s adjustedSelf-hosted, RTX 5090
Hopper (HopitAI, Qwen3.5-4B LoRA)#1567.50.12s raw, 0.39s adjustedSelf-hosted, RTX A6000
djev (Maisa, DiffusionGemma)#2064.20.05s raw, 0.25s adjustedSelf-hosted, H100
Laya (ModernBERT-large)#930.00.67s raw, 1.49s adjustedSelf-hosted, CPU

JevBench v1.5 is a new protocol, not an update of v1.4. Every system answers 1,624 decisions: 904 open decisions from public sources and 720 sealed decisions whose text stays private, one request at a time. The official score is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard; an axis below chance counts as zero, which zeroes the score. The sealed decisions make up half of intelligence, and a system whose open-minus-sealed gap exceeds the field median loses intelligence in proportion. Choice, noul and score requests carry equal weight. Benchmark Heaven also leads its page with a Capability ranking: the mean of intelligence and calibration among Jev-class systems, those within twice Jev’s cost and median latency. v1.5.4 lists 112 systems and ranks 106; releases since v1.5.4 add systems measured with the frozen method and change no earlier score. Self-hosted latency is adjusted by x2 plus 0.15s to approximate production load, so it is not comparable with the raw latency authors publish. Scores are not comparable with any v1.4 release.

Jev and Mercury Decide feature by feature

AttributeJev 1.13.0Mercury Decide
What it isHosted System One model from TypeSafe, version jev-1.13.0Hosted decision model from Inception, built on diffusion language modelling
Request formatstate plus noul, choice and score questionsThe same schema; works with the same Jev AI request
ProbabilitiesPer option, with a confidence valuePer option, with a confidence value, read from the model’s distribution
Context64k tokens per request32,768 tokens per request
Independent score72.0 intelligence in JevBench v1.5.4Not measured in JevBench v1.5.4; Inception reports leading JevBench v1.4
Agreement on our examplesReference122 of 133 answers (91.7%) match Jev
Speed0.65s median on the production APIUp to 14 decisions per second by Inception’s account; 0.67s median wall-clock in our run
Model name on Jev AIjev-latest, jev-preview or jev-1.13.0mercury-decide (Decisions API alias inception/mercury-decide)

What each one is better at

Where Jev wins

  • Twice the context: 64k tokens against 32k, so long tickets, transcripts and documents fit in one request.
  • An independent, current benchmark record in JevBench v1.5.4, including calibration.
  • Generally available rather than an early-access endpoint, with pinnable versions such as jev-1.13.0.
  • Most of our disagreements were on borderline yes/no and rubric scores, where Jev is the model these templates were written and checked against.

Where Mercury Decide wins

  • Diffusion decoding built for throughput: Inception quotes up to 14 decisions per second.
  • A drop-in alternative on the same schema, so you can switch with one field and keep every judge and template.
  • Close agreement with Jev on everyday tasks: 53 of 54 choices matched in our replay.
  • Slightly fewer counted input tokens on the same requests in our sample.

The analysis

Where Jev and Mercury Decide really differ, and why the bars above land where they do.

Inception released Mercury Decide on 30 September 2026 as a decision model rather than a chat model. It accepts the same request Jev does and returns probabilities read from the model’s own distribution instead of generated text, which is why output tokens cost nothing. Inception says it leads JevBench v1.4 and runs up to 14 decisions per second.

Benchmark Heaven’s current release, JevBench v1.5.4, does not list it, so there is no independent score to chart here. Instead we replayed every recorded Jev request on this site — 8 homepage demos and 33 use-case templates, 41 requests and 133 questions — on Mercury Decide and compared the answers. 91.7% agree: 53 of 54 choices, 37 of 42 yes/no answers on the same side of 0.5, and 32 of 37 scores at the same level.

Agreement is not accuracy. These examples have no ground-truth labels, so where the two differ we cannot say which is right; read the numbers as how interchangeable the models are on everyday triage, routing, moderation and matching tasks. On that measure they are close, with most disagreement on borderline yes/no and graded questions rather than on clear-cut choices.

What Mercury Decide actually is

Mercury Decide, Inception’s diffusion-based decision model, by Inception.

Mercury Decide comes from Inception, the company behind the Mercury diffusion language models. Like Mercury 2.5, it refines many positions in parallel over a few passes instead of producing one token at a time. For a decision model that matters less for the answer itself than for speed: the answer is a probability over a short set of labels, not a paragraph.

It supports the three System One question types: a choice among named options, a score on an ordered rubric, and a yes/no probability. Inception aims it at moderation, routing, fraud triage, agent tool selection, evaluation loops, grading, policy checks and escalation gates — the same territory as Jev.

It is served on OpenRouter as an early-access endpoint with a 32,768-token context window. Inception notes that real latency depends on request size, network conditions, provider load and concurrency, and that a public benchmark cannot establish accuracy or calibration on your own data.

How we ran the head-to-head

Inputs were the site’s published, fictional examples: each homepage demo and each use-case template, exactly as the playground sends them. Jev’s side is the recorded jev-1.13.0 answer already shown on those pages; Mercury’s side was run on 1 October 2026 against snapshot inception/mercury-decide-20260930, one request at a time. No request failed.

A choice agrees when both pick the same option. A yes/no answer agrees when both probabilities fall on the same side of 0.5; the mean gap between the two probabilities was 0.147. A score agrees when both round to the same rubric level; the mean gap was 0.148 of a level. The script is in the repository as scripts/bench-mercury-vs-jev.ts, so the run can be repeated.

Mercury’s median wall-clock time in this run was 0.67s, measured from one client and including the network hop to OpenRouter, so it is not comparable with lab latency. On the eight demos where Jev’s usage was recorded, Mercury counted 3,364 input tokens against Jev’s 3,749 for the same requests; the models tokenise differently.

So which should you use?

Pick Jev if

  • Inputs are long — a full transcript, contract or multi-message thread.
  • You need a version you can pin and an independently benchmarked model.
  • Calibrated yes/no and rubric scores on borderline cases matter more than raw throughput.

Pick Mercury Decide if

  • You run large volumes of short decisions and want a second fast model on the same schema.
  • You want a second opinion: send the same request to both and escalate when they disagree.
  • Inputs stay well under 32k tokens.

Try Jev on the cases you are actually arguing about

Because the request is identical, this is the rare comparison you can settle on your own data in minutes. Run your judge on Jev and on Mercury Decide side by side, look at the cases where they disagree, and keep whichever model is right more often on those. Where both agree, you can trust the answer more than either alone.

Jev vs Mercury Decide FAQ

Is Mercury Decide compatible with Jev requests?

Yes. It uses the same System One schema: a state plus typed noul, choice and score questions. On Jev AI, set model to mercury-decide and keep the rest of the request, including saved judges.

Is Mercury Decide better than Jev?

There is no independent current score yet: JevBench v1.5.4 does not include it. In our replay of 41 recorded requests the two agreed on 91.7% of 133 answers. Inception reports leading JevBench v1.4, an older protocol whose scores are not comparable with v1.5.

How long can the input be?

Mercury Decide accepts 32,768 tokens per request; Jev accepts 64k. Longer inputs to Mercury Decide are rejected without a charge.

Can I call Mercury Decide through the Jev AI API?

Yes. Use POST /api/v1/systemone with model mercury-decide, or the Decisions endpoint with inception/mercury-decide. The same API key, balance and billing rules apply as for Jev.

Other Jev comparisons

Every number on these pages is quoted from a published source and was read on September 30, 2026.

Sources

Mercury Decide and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on September 30, 2026 and can change.