On this site
Jev 1.13.0
TypeSafe · System One model
Third overall, fourth on Capability, and ahead on intelligence and every request type.
71.5JevBench v1.6.1 rank #3 · Capability #4
Jev comparison · JevBench v1.6.1 · updated October 8, 2026
The newest name in the top ten, and it trains nothing of its own.
On this site
TypeSafe · System One model
Third overall, fourth on Capability, and ahead on intelligence and every request type.
71.5JevBench v1.6.1 rank #3 · Capability #4
vs
Alternative
Amin Roudaki · Frozen Gemma 4 12B behind the /v1/systemone wire format
Tenth of 127 from a frozen base, and twelfth on Capability.
68.2JevBench v1.6.1 rank #10 · Capability #12
One benchmark measured all 165 systems under one method and ranked 127, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.
The top four ranked systems finish within 2.6 points of each other. The composite hides where systems actually differ, so the charts below break it apart. Benchmark Heaven’s headline Capability ranking, shown next to each score above, averages intelligence and calibration among Jev-class systems.
IntelligenceHow often it picks the right answer, half from sealed decisions
CalibrationWhether 0.8 really means about 80%
SpeedMeasured response time
Open300 decisions from public sources
Sealed1,200 private decisions nobody could tune for
ChoicePick one of the options
NoulProbability that a statement is true
ScorePlace the case on a graded scale
The sealed half, highlighted, is where results cannot have been tuned: nobody outside the benchmark has seen those decisions. Request types are scored above chance, so 0 means guessing and a value below 0 means worse than guessing.
One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.
The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and decisio are highlighted.
| Model | Rank | Score | Measured latency | Runs on |
|---|---|---|---|---|
| Sage 1.3.0 (Levanto Labs, hosted API) | #1 | 74.0 | 0.15s median / 0.29s p95 | Hosted API |
| H2O-Lightning-4B v1.1 (H2O.ai, Qwen3.5-4B) | #2 | 72.5 | 0.03s raw, 0.21s adjusted | Self-hosted GPU |
| Jev 1.13.0 (TypeSafe) | #3 | 71.5 | 0.24s median / 0.30s p95 | Hosted API |
| Quyet-1.0-Large (Chinh Nguyen, Gemma 4 31B decoder) | #4 | 71.4 | 0.12s raw, 0.38s adjusted | Self-hosted, H100 |
| wity-1 (Wity, closed API, reasoning auto) | #5 | 70.8 | 1.57s median / 4.21s p95 | Hosted API |
| torchcast-decision-12b (Torchcast AI, Gemma 4 12B) | #6 | 69.9 | 0.04s raw, 0.22s adjusted | Self-hosted, H100 |
| Winnow-12B Q8 (EldanRing, Gemma 4 12B) | #7 | 68.9 | 0.12s raw, 0.38s adjusted | Self-hosted, RTX 6000 |
| deck-31B (krishna765, frozen Gemma 4 31B) | #8 | 68.7 | 0.13s raw, 0.40s adjusted | Self-hosted GPU |
| Cygnet (blockbrain, frozen Gemma 4 12B) | #9 | 68.6 | 0.04s raw, 0.22s adjusted | Self-hosted, RTX PRO 6000 |
| decisio v0.8.0 (Amin Roudaki, frozen Gemma 4 12B) | #10 | 68.2 | 0.10s raw, 0.35s adjusted | Self-hosted GPU |
| Jev-Omni (akhilaaa3, Gemma 4 12B) | #12 | 67.7 | 0.15s raw, 0.46s adjusted | Self-hosted, RTX 6000 |
| Xor 26B-A4B (Juspay, Gemma 4 26B-A4B) | #13 | 67.4 | 0.06s raw, 0.27s adjusted | Self-hosted GPU |
| Plumb-4B (crh225, JevK5 v0.2 + LoRA) | #23 | 48.3 | 0.02s raw, 0.19s adjusted | Self-hosted, H100 |
| decider-4b v2 (Mapika, Qwen3.5-4B) | #31 | 41.2 | 0.02s raw, 0.20s adjusted | Self-hosted, RTX 5090 |
| JevK5 v0.3 (allebee, Qwen3.5-4B) | #33 | 37.4 | 0.02s raw, 0.19s adjusted | Self-hosted, H100 |
| Hopper (HopitAI, Qwen3.5-4B LoRA) | #44 | 28.9 | 0.06s raw, 0.26s adjusted | Self-hosted, RTX A6000 |
| Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA) | #45 | 28.7 | 0.07s raw, 0.29s adjusted | Self-hosted, RTX 5090 |
| SemIf (Qwen3.5-4B) | #72 | 15.5 | 0.05s raw, 0.25s adjusted | Self-hosted, RTX 5090 |
| Laya (ModernBERT-large) | #119 | 0.0 | 0.83s raw, 1.80s adjusted | Self-hosted, CPU |
JevBench v1.6 is a full re-measure on a new decision pool, not an update of v1.5. Every system answers 1,500 decisions: 300 open decisions from public sources and 1,200 sealed decisions whose text stays private, one request at a time. v1.6.1 put the hosted APIs on that same full set instead of the 600-item subset v1.6.0 gave them, so their accuracy axes are no longer equated onto the self-hosted scale. The official score is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard; an axis below chance counts as zero, which zeroes the score. The sealed decisions make up half of intelligence, and a system whose open-minus-sealed gap exceeds the field median loses intelligence in proportion. Choice, noul and score requests carry equal weight. v1.6.1 also prices every token-priced system on one common item set — all 1,500 decisions except the 23 that fall outside Jev’s accepted input range — so a system that answered those long items no longer pays for their tokens while one that refused them does not. Refusals still count against intelligence. Benchmark Heaven also leads its page with a Capability ranking: the mean of intelligence and calibration among Jev-class systems, those inside a budget frozen at twice Jev’s v1.5 cost and median latency. v1.6.1 lists 165 systems and ranks 127; the rest were not re-measured, and 13 of them keep a separately dated v1.5.x score. Self-hosted latency is adjusted by x2 plus 0.15s to approximate production load, so it is not comparable with the raw latency authors publish. Scores are not comparable with any v1.5 or v1.4 release.
| Attribute | Jev 1.13.0 | decisio |
|---|---|---|
| What it is | Hosted System One model from TypeSafe, version jev-1.13.0 | Apache-2.0 serving layer over a frozen base; measured on google/gemma-4-12B-it, bf16 |
| Licence | Proprietary, hosted | Apache-2.0 software; Gemma 4 weights under Google’s model terms |
| Request format | The original /v1/systemone | The same /v1/systemone, plus /v1/tasks and an abstention route of its own |
| Question types | Noul, choice, score — many per call, answered in parallel | The same three, many per request against one shared prefill |
| Options per choice | Up to 255 | Up to 255 |
| Supported inputs | Text only | Text, plus images through a second engine on the same card |
| Learning from your examples | None — the model is fixed | Registers a task from labelled examples: a calibration, and an intent head for long option lists |
| Base model | Undisclosed, post-trained by TypeSafe | Yours to choose: Qwen3.6-35B-A3B, Gemma 4 12B or Gemma 4 31B, each frozen |
| Hardware | None — it is an API call | One NVIDIA card under vLLM; also Apple silicon, Ollama or CPU for development |
Where Jev and decisio really differ, and why the bars above land where they do.
Three systems in the current top ten train nothing at all — deck-31B at #8, Cygnet at #9 and decisio at #10 — and decisio is the newest of the three and by far the most featureful. It is #10 of 127 at 68.2 against Jev’s 71.5 at #3, and unlike the other two it is not a thin shim. It speaks the same /v1/systemone wire format as Jev, accepts up to 255 options per choice exactly as Jev does, prefills a state once and answers many questions against it, and adds things Jev has no equivalent for: registering a recurring question from labelled examples, an opt-in abstention threshold, and images in the state through a second engine.
On the answers Jev is ahead everywhere. Intelligence is 63.6 against 52.3, and 61.6 against 51.6 on the sealed half. By request type Jev leads on choice questions, 77.4 against 69.1, on yes/no, 50.3 against 42.2, and most heavily on score questions, 63.2 against 45.7 — a 17.5-point gap that is the single biggest reason for the difference in the composite.
Where it closes in is the two axes that are not about being right. Calibration is 89.2 against Jev’s 90.6, which is close for a model nobody trained, and its cost score is the better of the two, 59.3 against 54.7. Jev is faster on the benchmark’s measure, 91.5 against 87.0. On Benchmark Heaven’s headline Capability ranking, which averages intelligence and calibration, Jev is fourth at 77.1 and decisio twelfth at 70.7. Jev also stays above it under all three of the benchmark’s orders — third, fourth and third against tenth, thirteenth and fourteenth — so no reweighting closes the gap.
decisio, the decision serving layer over frozen open bases, by Amin Roudaki.
decisio is an Apache-2.0 Python package from Amin Roudaki, first published on 30 September 2026. It implements TypeSafe’s published System One wire format and is not affiliated with TypeSafe. Typed questions about a piece of text go in; a probability for every option comes out, each question from one forward pass with nothing generated.
The base model is a choice, not part of the project. Three ship as profiles behind the same routes and features: Qwen3.6-35B-A3B FP8 as the default, google/gemma-4-12B-it in bf16, and google/gemma-4-31B-it quantised to FP8 on load. Each is an official checkpoint at a pinned revision with no adapter and no fine-tuning, and each profile carries its own measured settings. JevBench measured the Gemma 4 12B profile, whose weights take 22.8 GiB in memory and whose fitted temperature is 3.592 for every question type.
Beyond answering, it registers tasks: POST /v1/tasks learns one recurring question from labelled examples, fitting a calibration and, for long option lists, an intent head. Abstention is opt-in per task against a declared “can’t tell” option. Images in the state are served by a second engine on the same card. A state read once is reused from the prefix cache, so a second, different question on the same document answers in about 36ms on the Gemma 4 12B base against about 260ms for the first read.
It runs on one NVIDIA card under vLLM, in Docker, on Apple silicon through MLX, through Ollama on a 16 GB laptop for the smallest Gemma tag, or on CPU as a development stand-in. The software is Apache-2.0; the Gemma weights carry Google’s model terms.
Both v0.8.0 and v0.9.0 of the Gemma 4 12B profile are on the board, at #10 with 68.2 and #11 with 67.8. Their intelligence is the same to a tenth, 52.3 either way: v0.9.0 answers no better and no worse.
The difference is on the other axes. Calibration falls from 89.2 to 87.6 and the speed score from 87.0 to 85.9, and the harmonic mean turns that into the four tenths between them. The change v0.9.0 describes for this base is when a state’s boundary is registered — after the answer rather than before — which the author measured as a faster first read on an idle queue. Whatever it gained there, the benchmark’s adjusted median did not improve.
This page quotes v0.8.0 throughout, because it is the higher-placed of the two and the one the chart shows.
On yes/no questions decisio scores 42.2 above chance against Jev’s 50.3, its widest relative shortfall after score questions. The underlying behaviour is not random guessing: the benchmark records it as committing to an answer on 81.7% of the yes/no items, and being right on 90.8% of the ones it committed to.
The rest is hedging near the middle of the range, which a chance-corrected score treats as close to no answer at all. decisio has a flag for exactly this, reporting a probability inside the no-answer band at the band’s edge, which changes the calibration rather than the answer; the configuration JevBench measured did not use it. If your code reads the probability rather than just the top label, that distinction is the one to test on your own cases.
Its answers are also almost always well formed: the benchmark records invalid-response rates of 1.7% on choice questions, 1.1% on yes/no and 1.6% on score.
decisio is the strongest argument on the board that the serving layer, not the training run, is what most decision work is missing — and it is still behind Jev on the answers. Both are worth measuring on your own decisions rather than on either ranking. Run them on Jev here with no setup and keep the answers.
Not on the answers. JevBench v1.6.1 ranks Jev #3 at 71.5 and decisio #10 at 68.2, and Jev leads on intelligence, 63.6 against 52.3, on the sealed half, 61.6 against 51.6, and on all three request types. decisio has the better cost score and nearly matches Jev on calibration, 89.2 against 90.6.
No. Every base it serves is an official checkpoint at a pinned revision with no adapter and no fine-tuning. The fitted values are calibration temperatures, 3.592 for every question type on the Gemma 4 12B profile JevBench measured.
google/gemma-4-12B-it in bf16, which is the profile that placed #10. decisio’s default base is Qwen3.6-35B-A3B, and it also serves Gemma 4 31B, which the author measures as the strongest of the three; neither is ranked on the open-weights board.
JevBench ran the Gemma 4 12B profile on one RTX 6000 Ada with 48 GB. That base takes 22.8 GiB in memory and does not fit a 32 GB card under vLLM. It also runs on Apple silicon through MLX, through Ollama on a 16 GB laptop for the smallest Gemma tag, and on CPU as a development stand-in.
It implements TypeSafe’s published System One wire format, so requests in that shape are what it accepts. It is an independent project and not affiliated with TypeSafe, and it adds routes of its own for task registration and abstention.
Two versions of the Gemma 4 12B profile were measured: v0.8.0 at #10 with 68.2 and v0.9.0 at #11 with 67.8. Their intelligence is identical at 52.3; v0.9.0 scores lower on calibration and speed. This page quotes v0.8.0, the higher-placed row.
Every number on these pages is quoted from a published source and was read on October 8, 2026.
decisio and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 8, 2026 and can change.