On this site
Jev 1.13.0
TypeSafe · System One model
The original System One model, with 64k of context and a public benchmark record.
72.1JevBench v1.5.4 rank #3 · Capability #1
Jev comparison · JevBench v1.5.4 · updated September 30, 2026
Two hosted decision models that take the same request.
On this site
TypeSafe · System One model
The original System One model, with 64k of context and a public benchmark record.
72.1JevBench v1.5.4 rank #3 · Capability #1
vs
Alternative
Inception · Hosted diffusion decision model
Same schema, diffusion decoding, and up to 14 decisions per second by Inception’s count.
Not measured in JevBench v1.5.4
Both models run on your Jev AI account with the same request: switch the model in the playground, or set model to mercury-decide in the API. Mercury Decide model guide →
One benchmark measured all 112 systems under one method and ranked 106, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.
The top four ranked systems finish within 1.8 points of each other. The composite hides where systems actually differ, so the charts below break it apart. Benchmark Heaven’s headline Capability ranking, shown next to each score above, averages intelligence and calibration among Jev-class systems.
Mercury Decide was not measured in JevBench v1.5.4, so there is nothing to chart against Jev here. The analysis below uses the figures its authors publish.
One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.
The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and Mercury Decide are highlighted.
| Model | Rank | Score | Measured latency | Runs on |
|---|---|---|---|---|
| Cygnet (blockbrain, frozen Gemma 4 12B) | #1 | 73.7 | 0.04s raw, 0.23s adjusted | Self-hosted, RTX PRO 6000 |
| Winnow-12B Q8 (EldanRing, Gemma 4 12B) | #2 | 73.2 | 0.10s raw, 0.34s adjusted | Self-hosted, RTX 6000 |
| Jev 1.13.0 (TypeSafe) | #3 | 72.1 | 0.62s median / 0.67s p95 | Hosted API |
| JevK5 v0.3 (allebee, Qwen3.5-4B) | #4 | 71.9 | 0.02s raw, 0.18s adjusted | Self-hosted, H100 |
| Plumb-4B (crh225, JevK5 v0.2 + LoRA) | #5 | 71.6 | 0.02s raw, 0.18s adjusted | Self-hosted, H100 |
| Jev-Omni (akhilaaa3, Gemma 4 12B) | #6 | 71.5 | 0.11s raw, 0.38s adjusted | Self-hosted, RTX 6000 |
| decider-4b v2 (Mapika, Qwen3.5-4B) | #7 | 71.3 | 0.03s raw, 0.20s adjusted | Self-hosted, RTX 5090 |
| Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA) | #9 | 70.4 | 0.04s raw, 0.23s adjusted | Self-hosted, RTX 5090 |
| SemIf (Qwen3.5-4B) | #12 | 68.7 | 0.04s raw, 0.23s adjusted | Self-hosted, RTX 5090 |
| Hopper (HopitAI, Qwen3.5-4B LoRA) | #15 | 67.5 | 0.12s raw, 0.39s adjusted | Self-hosted, RTX A6000 |
| djev (Maisa, DiffusionGemma) | #20 | 64.2 | 0.05s raw, 0.25s adjusted | Self-hosted, H100 |
| Laya (ModernBERT-large) | #93 | 0.0 | 0.67s raw, 1.49s adjusted | Self-hosted, CPU |
JevBench v1.5 is a new protocol, not an update of v1.4. Every system answers 1,624 decisions: 904 open decisions from public sources and 720 sealed decisions whose text stays private, one request at a time. The official score is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard; an axis below chance counts as zero, which zeroes the score. The sealed decisions make up half of intelligence, and a system whose open-minus-sealed gap exceeds the field median loses intelligence in proportion. Choice, noul and score requests carry equal weight. Benchmark Heaven also leads its page with a Capability ranking: the mean of intelligence and calibration among Jev-class systems, those within twice Jev’s cost and median latency. v1.5.4 lists 112 systems and ranks 106; releases since v1.5.4 add systems measured with the frozen method and change no earlier score. Self-hosted latency is adjusted by x2 plus 0.15s to approximate production load, so it is not comparable with the raw latency authors publish. Scores are not comparable with any v1.4 release.
| Attribute | Jev 1.13.0 | Mercury Decide |
|---|---|---|
| What it is | Hosted System One model from TypeSafe, version jev-1.13.0 | Hosted decision model from Inception, built on diffusion language modelling |
| Request format | state plus noul, choice and score questions | The same schema; works with the same Jev AI request |
| Probabilities | Per option, with a confidence value | Per option, with a confidence value, read from the model’s distribution |
| Context | 64k tokens per request | 32,768 tokens per request |
| Independent score | 72.0 intelligence in JevBench v1.5.4 | Not measured in JevBench v1.5.4; Inception reports leading JevBench v1.4 |
| Agreement on our examples | Reference | 122 of 133 answers (91.7%) match Jev |
| Speed | 0.65s median on the production API | Up to 14 decisions per second by Inception’s account; 0.67s median wall-clock in our run |
| Model name on Jev AI | jev-latest, jev-preview or jev-1.13.0 | mercury-decide (Decisions API alias inception/mercury-decide) |
Where Jev and Mercury Decide really differ, and why the bars above land where they do.
Inception released Mercury Decide on 30 September 2026 as a decision model rather than a chat model. It accepts the same request Jev does and returns probabilities read from the model’s own distribution instead of generated text, which is why output tokens cost nothing. Inception says it leads JevBench v1.4 and runs up to 14 decisions per second.
Benchmark Heaven’s current release, JevBench v1.5.4, does not list it, so there is no independent score to chart here. Instead we replayed every recorded Jev request on this site — 8 homepage demos and 33 use-case templates, 41 requests and 133 questions — on Mercury Decide and compared the answers. 91.7% agree: 53 of 54 choices, 37 of 42 yes/no answers on the same side of 0.5, and 32 of 37 scores at the same level.
Agreement is not accuracy. These examples have no ground-truth labels, so where the two differ we cannot say which is right; read the numbers as how interchangeable the models are on everyday triage, routing, moderation and matching tasks. On that measure they are close, with most disagreement on borderline yes/no and graded questions rather than on clear-cut choices.
Mercury Decide, Inception’s diffusion-based decision model, by Inception.
Mercury Decide comes from Inception, the company behind the Mercury diffusion language models. Like Mercury 2.5, it refines many positions in parallel over a few passes instead of producing one token at a time. For a decision model that matters less for the answer itself than for speed: the answer is a probability over a short set of labels, not a paragraph.
It supports the three System One question types: a choice among named options, a score on an ordered rubric, and a yes/no probability. Inception aims it at moderation, routing, fraud triage, agent tool selection, evaluation loops, grading, policy checks and escalation gates — the same territory as Jev.
It is served on OpenRouter as an early-access endpoint with a 32,768-token context window. Inception notes that real latency depends on request size, network conditions, provider load and concurrency, and that a public benchmark cannot establish accuracy or calibration on your own data.
Inputs were the site’s published, fictional examples: each homepage demo and each use-case template, exactly as the playground sends them. Jev’s side is the recorded jev-1.13.0 answer already shown on those pages; Mercury’s side was run on 1 October 2026 against snapshot inception/mercury-decide-20260930, one request at a time. No request failed.
A choice agrees when both pick the same option. A yes/no answer agrees when both probabilities fall on the same side of 0.5; the mean gap between the two probabilities was 0.147. A score agrees when both round to the same rubric level; the mean gap was 0.148 of a level. The script is in the repository as scripts/bench-mercury-vs-jev.ts, so the run can be repeated.
Mercury’s median wall-clock time in this run was 0.67s, measured from one client and including the network hop to OpenRouter, so it is not comparable with lab latency. On the eight demos where Jev’s usage was recorded, Mercury counted 3,364 input tokens against Jev’s 3,749 for the same requests; the models tokenise differently.
Because the request is identical, this is the rare comparison you can settle on your own data in minutes. Run your judge on Jev and on Mercury Decide side by side, look at the cases where they disagree, and keep whichever model is right more often on those. Where both agree, you can trust the answer more than either alone.
Yes. It uses the same System One schema: a state plus typed noul, choice and score questions. On Jev AI, set model to mercury-decide and keep the rest of the request, including saved judges.
There is no independent current score yet: JevBench v1.5.4 does not include it. In our replay of 41 recorded requests the two agreed on 91.7% of 133 answers. Inception reports leading JevBench v1.4, an older protocol whose scores are not comparable with v1.5.
Mercury Decide accepts 32,768 tokens per request; Jev accepts 64k. Longer inputs to Mercury Decide are rejected without a charge.
Yes. Use POST /api/v1/systemone with model mercury-decide, or the Decisions endpoint with inception/mercury-decide. The same API key, balance and billing rules apply as for Jev.
Every number on these pages is quoted from a published source and was read on September 30, 2026.
Mercury Decide and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on September 30, 2026 and can change.