Jev AI

Jev comparison · JevBench v1.6.1 · updated October 8, 2026

Jev vs decisio

The newest name in the top ten, and it trains nothing of its own.

On this site

Jev 1.13.0

TypeSafe · System One model

Third overall, fourth on Capability, and ahead on intelligence and every request type.

71.5JevBench v1.6.1 rank #3 · Capability #4

vs

Alternative

decisio

Amin Roudaki · Frozen Gemma 4 12B behind the /v1/systemone wire format

Tenth of 127 from a frozen base, and twelfth on Capability.

68.2JevBench v1.6.1 rank #10 · Capability #12

  1. JevHosted API, generally availableWhere it runsdecisioSelf-hosted on vLLM, Docker, Apple silicon or Ollama
  2. JevPost-trained decision modelTrainingdecisioNone — official Gemma 4 12B at a pinned revision, no adapter
  3. JevUp to 255Options per choicedecisioUp to 255
  4. Jev90.6 on JevBenchCalibrationdecisio89.2 on JevBench
  5. Jev63.6 on JevBenchIntelligencedecisio52.3 on JevBench
  6. Jev61.6Sealed decisionsdecisio51.6
  7. Jev63.2 above chanceScore questionsdecisio45.7 above chance
  8. Jev91.5 · 0.24s median over the networkSpeeddecisio87.0 · 0.10s raw on the benchmark’s GPU

JevBench v1.6.1, measured the same way

One benchmark measured all 165 systems under one method and ranked 127, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.

Composite score

#1 Sage 1.3.074.0
#2 H2O-Lightning-4B72.5
#3 Jev71.5
#4 Quyet-1.0-Large71.4
#5 wity-170.8
#6 torchcast-12b69.9
#7 Winnow68.9
#8 deck-31B68.7
#9 Cygnet68.6
#10 decisio v0.8.068.2
#12 Jev-Omni67.7
#13 Xor 26B-A4B67.4
#23 Plumb-4B48.3
#31 decider-4b41.2
#33 JevK5 v0.337.4
#44 Hopper28.9
#45 Imajev-4B28.7
#72 SemIf15.5
#119 Laya0.0

The top four ranked systems finish within 2.6 points of each other. The composite hides where systems actually differ, so the charts below break it apart. Benchmark Heaven’s headline Capability ranking, shown next to each score above, averages intelligence and calibration among Jev-class systems.

Capability, higher is better

IntelligenceHow often it picks the right answer, half from sealed decisions

Jev63.6
decisio52.3

CalibrationWhether 0.8 really means about 80%

Jev90.6
decisio89.2

SpeedMeasured response time

Jev91.5
decisio87.0

Intelligence on open and sealed decisions

Open300 decisions from public sources

Jev65.6
decisio53.1

Sealed1,200 private decisions nobody could tune for

Jev61.6
decisio51.6

By request type, chance-corrected

ChoicePick one of the options

Jev77.4
decisio69.1

NoulProbability that a statement is true

Jev50.3
decisio42.2

ScorePlace the case on a graded scale

Jev63.2
decisio45.7

The sealed half, highlighted, is where results cannot have been tuned: nobody outside the benchmark has seen those decisions. Request types are scored above chance, so 0 means guessing and a value below 0 means worse than guessing.

One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.

Every system on the board

The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and decisio are highlighted.

ModelRankScoreMeasured latencyRuns on
Sage 1.3.0 (Levanto Labs, hosted API)#174.00.15s median / 0.29s p95Hosted API
H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B)#272.50.03s raw, 0.21s adjustedSelf-hosted GPU
Jev 1.13.0 (TypeSafe)#371.50.24s median / 0.30s p95Hosted API
Quyet-1.0-Large (Chinh Nguyen, Gemma 4 31B decoder)#471.40.12s raw, 0.38s adjustedSelf-hosted, H100
wity-1 (Wity, closed API, reasoning auto)#570.81.57s median / 4.21s p95Hosted API
torchcast-decision-12b (Torchcast AI, Gemma 4 12B)#669.90.04s raw, 0.22s adjustedSelf-hosted, H100
Winnow-12B Q8 (EldanRing, Gemma 4 12B)#768.90.12s raw, 0.38s adjustedSelf-hosted, RTX 6000
deck-31B (krishna765, frozen Gemma 4 31B)#868.70.13s raw, 0.40s adjustedSelf-hosted GPU
Cygnet (blockbrain, frozen Gemma 4 12B)#968.60.04s raw, 0.22s adjustedSelf-hosted, RTX PRO 6000
decisio v0.8.0 (Amin Roudaki, frozen Gemma 4 12B)#1068.20.10s raw, 0.35s adjustedSelf-hosted GPU
Jev-Omni (akhilaaa3, Gemma 4 12B)#1267.70.15s raw, 0.46s adjustedSelf-hosted, RTX 6000
Xor 26B-A4B (Juspay, Gemma 4 26B-A4B)#1367.40.06s raw, 0.27s adjustedSelf-hosted GPU
Plumb-4B (crh225, JevK5 v0.2 + LoRA)#2348.30.02s raw, 0.19s adjustedSelf-hosted, H100
decider-4b v2 (Mapika, Qwen3.5-4B)#3141.20.02s raw, 0.20s adjustedSelf-hosted, RTX 5090
JevK5 v0.3 (allebee, Qwen3.5-4B)#3337.40.02s raw, 0.19s adjustedSelf-hosted, H100
Hopper (HopitAI, Qwen3.5-4B LoRA)#4428.90.06s raw, 0.26s adjustedSelf-hosted, RTX A6000
Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA)#4528.70.07s raw, 0.29s adjustedSelf-hosted, RTX 5090
SemIf (Qwen3.5-4B)#7215.50.05s raw, 0.25s adjustedSelf-hosted, RTX 5090
Laya (ModernBERT-large)#1190.00.83s raw, 1.80s adjustedSelf-hosted, CPU

JevBench v1.6 is a full re-measure on a new decision pool, not an update of v1.5. Every system answers 1,500 decisions: 300 open decisions from public sources and 1,200 sealed decisions whose text stays private, one request at a time. v1.6.1 put the hosted APIs on that same full set instead of the 600-item subset v1.6.0 gave them, so their accuracy axes are no longer equated onto the self-hosted scale. The official score is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard; an axis below chance counts as zero, which zeroes the score. The sealed decisions make up half of intelligence, and a system whose open-minus-sealed gap exceeds the field median loses intelligence in proportion. Choice, noul and score requests carry equal weight. v1.6.1 also prices every token-priced system on one common item set — all 1,500 decisions except the 23 that fall outside Jev’s accepted input range — so a system that answered those long items no longer pays for their tokens while one that refused them does not. Refusals still count against intelligence. Benchmark Heaven also leads its page with a Capability ranking: the mean of intelligence and calibration among Jev-class systems, those inside a budget frozen at twice Jev’s v1.5 cost and median latency. v1.6.1 lists 165 systems and ranks 127; the rest were not re-measured, and 13 of them keep a separately dated v1.5.x score. Self-hosted latency is adjusted by x2 plus 0.15s to approximate production load, so it is not comparable with the raw latency authors publish. Scores are not comparable with any v1.5 or v1.4 release.

Jev and decisio feature by feature

AttributeJev 1.13.0decisio
What it isHosted System One model from TypeSafe, version jev-1.13.0Apache-2.0 serving layer over a frozen base; measured on google/gemma-4-12B-it, bf16
LicenceProprietary, hostedApache-2.0 software; Gemma 4 weights under Google’s model terms
Request formatThe original /v1/systemoneThe same /v1/systemone, plus /v1/tasks and an abstention route of its own
Question typesNoul, choice, score — many per call, answered in parallelThe same three, many per request against one shared prefill
Options per choiceUp to 255Up to 255
Supported inputsText onlyText, plus images through a second engine on the same card
Learning from your examplesNone — the model is fixedRegisters a task from labelled examples: a calibration, and an intent head for long option lists
Base modelUndisclosed, post-trained by TypeSafeYours to choose: Qwen3.6-35B-A3B, Gemma 4 12B or Gemma 4 31B, each frozen
HardwareNone — it is an API callOne NVIDIA card under vLLM; also Apple silicon, Ollama or CPU for development

What each one is better at

Where Jev wins

  • Intelligence 63.6 against 52.3, and 61.6 against 51.6 on the sealed half.
  • Ahead on all three request types: choice 77.4 against 69.1, yes/no 50.3 against 42.2 and score 63.2 against 45.7.
  • Fourth on the Capability ranking at 77.1 against twelfth at 70.7, and above decisio under all three published orders.
  • Faster on the benchmark’s measure, 91.5 against 87.0, with nothing to serve or keep running.
  • A released, versioned product rather than a package first published on 30 September 2026.

Where decisio wins

  • #10 of 127 in JevBench v1.6.1 at 68.2, from a base nobody fine-tuned, with only deck-31B at #8 and Cygnet at #9 above it among the untrained systems.
  • Up to 255 options per choice, the same as Jev, and many questions answered against one shared prefill.
  • Calibration 89.2 against Jev’s 90.6, and the better cost score of the two, 59.3 against 54.7.
  • Things Jev has no equivalent for: task registration from labelled examples, opt-in abstention and images in the state.
  • The base is a flag. Three frozen bases are served behind identical routes, and swapping one changes no application code.

The analysis

Where Jev and decisio really differ, and why the bars above land where they do.

Three systems in the current top ten train nothing at all — deck-31B at #8, Cygnet at #9 and decisio at #10 — and decisio is the newest of the three and by far the most featureful. It is #10 of 127 at 68.2 against Jev’s 71.5 at #3, and unlike the other two it is not a thin shim. It speaks the same /v1/systemone wire format as Jev, accepts up to 255 options per choice exactly as Jev does, prefills a state once and answers many questions against it, and adds things Jev has no equivalent for: registering a recurring question from labelled examples, an opt-in abstention threshold, and images in the state through a second engine.

On the answers Jev is ahead everywhere. Intelligence is 63.6 against 52.3, and 61.6 against 51.6 on the sealed half. By request type Jev leads on choice questions, 77.4 against 69.1, on yes/no, 50.3 against 42.2, and most heavily on score questions, 63.2 against 45.7 — a 17.5-point gap that is the single biggest reason for the difference in the composite.

Where it closes in is the two axes that are not about being right. Calibration is 89.2 against Jev’s 90.6, which is close for a model nobody trained, and its cost score is the better of the two, 59.3 against 54.7. Jev is faster on the benchmark’s measure, 91.5 against 87.0. On Benchmark Heaven’s headline Capability ranking, which averages intelligence and calibration, Jev is fourth at 77.1 and decisio twelfth at 70.7. Jev also stays above it under all three of the benchmark’s orders — third, fourth and third against tenth, thirteenth and fourteenth — so no reweighting closes the gap.

What decisio actually is

decisio, the decision serving layer over frozen open bases, by Amin Roudaki.

decisio is an Apache-2.0 Python package from Amin Roudaki, first published on 30 September 2026. It implements TypeSafe’s published System One wire format and is not affiliated with TypeSafe. Typed questions about a piece of text go in; a probability for every option comes out, each question from one forward pass with nothing generated.

The base model is a choice, not part of the project. Three ship as profiles behind the same routes and features: Qwen3.6-35B-A3B FP8 as the default, google/gemma-4-12B-it in bf16, and google/gemma-4-31B-it quantised to FP8 on load. Each is an official checkpoint at a pinned revision with no adapter and no fine-tuning, and each profile carries its own measured settings. JevBench measured the Gemma 4 12B profile, whose weights take 22.8 GiB in memory and whose fitted temperature is 3.592 for every question type.

Beyond answering, it registers tasks: POST /v1/tasks learns one recurring question from labelled examples, fitting a calibration and, for long option lists, an intent head. Abstention is opt-in per task against a declared “can’t tell” option. Images in the state are served by a second engine on the same card. A state read once is reused from the prefix cache, so a second, different question on the same document answers in about 36ms on the Gemma 4 12B base against about 260ms for the first read.

It runs on one NVIDIA card under vLLM, in Docker, on Apple silicon through MLX, through Ollama on a 16 GB laptop for the smallest Gemma tag, or on CPU as a development stand-in. The software is Apache-2.0; the Gemma weights carry Google’s model terms.

Two decisio versions are ranked, and the newer one is lower

Both v0.8.0 and v0.9.0 of the Gemma 4 12B profile are on the board, at #10 with 68.2 and #11 with 67.8. Their intelligence is the same to a tenth, 52.3 either way: v0.9.0 answers no better and no worse.

The difference is on the other axes. Calibration falls from 89.2 to 87.6 and the speed score from 87.0 to 85.9, and the harmonic mean turns that into the four tenths between them. The change v0.9.0 describes for this base is when a state’s boundary is registered — after the answer rather than before — which the author measured as a faster first read on an idle queue. Whatever it gained there, the benchmark’s adjusted median did not improve.

This page quotes v0.8.0 throughout, because it is the higher-placed of the two and the one the chart shows.

Why its yes/no score is its weakest axis

On yes/no questions decisio scores 42.2 above chance against Jev’s 50.3, its widest relative shortfall after score questions. The underlying behaviour is not random guessing: the benchmark records it as committing to an answer on 81.7% of the yes/no items, and being right on 90.8% of the ones it committed to.

The rest is hedging near the middle of the range, which a chance-corrected score treats as close to no answer at all. decisio has a flag for exactly this, reporting a probability inside the no-answer band at the band’s edge, which changes the calibration rather than the answer; the configuration JevBench measured did not use it. If your code reads the probability rather than just the top label, that distinction is the one to test on your own cases.

Its answers are also almost always well formed: the benchmark records invalid-response rates of 1.7% on choice questions, 1.1% on yes/no and 1.6% on score.

So which should you use?

Pick Jev if

  • Accuracy is what you are buying: Jev leads on intelligence, the sealed half and all three request types.
  • Score questions matter — the gap there is the widest of any axis, 63.2 against 45.7.
  • You would rather call an API than operate a GPU, a server and a base-model choice.
  • You want one versioned model with published limits, not a configuration you have to reproduce.

Pick decisio if

  • You want to keep the decision layer and the weights under your own control, on your own card, under Apache-2.0.
  • Your questions recur and you have labelled examples: registering a task is a lever Jev does not offer.
  • You need images in the state, or you want the option of moving to a stronger base without touching your code.
  • Volume is high and routine, and a better cost score matters more than the accuracy gap.

Try Jev on the cases you are actually arguing about

decisio is the strongest argument on the board that the serving layer, not the training run, is what most decision work is missing — and it is still behind Jev on the answers. Both are worth measuring on your own decisions rather than on either ranking. Run them on Jev here with no setup and keep the answers.

Jev vs decisio FAQ

Is decisio better than Jev?

Not on the answers. JevBench v1.6.1 ranks Jev #3 at 71.5 and decisio #10 at 68.2, and Jev leads on intelligence, 63.6 against 52.3, on the sealed half, 61.6 against 51.6, and on all three request types. decisio has the better cost score and nearly matches Jev on calibration, 89.2 against 90.6.

Is decisio fine-tuned?

No. Every base it serves is an official checkpoint at a pinned revision with no adapter and no fine-tuning. The fitted values are calibration temperatures, 3.592 for every question type on the Gemma 4 12B profile JevBench measured.

Which base did JevBench measure?

google/gemma-4-12B-it in bf16, which is the profile that placed #10. decisio’s default base is Qwen3.6-35B-A3B, and it also serves Gemma 4 31B, which the author measures as the strongest of the three; neither is ranked on the open-weights board.

What hardware does decisio need?

JevBench ran the Gemma 4 12B profile on one RTX 6000 Ada with 48 GB. That base takes 22.8 GiB in memory and does not fit a 32 GB card under vLLM. It also runs on Apple silicon through MLX, through Ollama on a 16 GB laptop for the smallest Gemma tag, and on CPU as a development stand-in.

Can I point the TypeSafe SDK at decisio?

It implements TypeSafe’s published System One wire format, so requests in that shape are what it accepts. It is an independent project and not affiliated with TypeSafe, and it adds routes of its own for task registration and abstention.

Why are there two decisio rows on the board?

Two versions of the Gemma 4 12B profile were measured: v0.8.0 at #10 with 68.2 and v0.9.0 at #11 with 67.8. Their intelligence is identical at 52.3; v0.9.0 scores lower on calibration and speed. This page quotes v0.8.0, the higher-placed row.

Other Jev comparisons

Every number on these pages is quoted from a published source and was read on October 8, 2026.

Sources

decisio and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 8, 2026 and can change.