Jev AI

Jev comparison · JevBench v1.4.2 · updated September 26, 2026

Jev vs decider-4b

The new first place on JevBench, and what that first place is made of.

On this site

Jev 1.13.0

TypeSafe · System One model

Second overall, and still first on intelligence, hard cases and the sealed set.

63.3JevBench v1.4.2 rank #2

vs

Alternative

decider-4b

Mapika · Open Qwen3.5-4B System One rebuild

First of 89 on the composite, from 8.4 GB of open weights.

64.1JevBench v1.4.2 rank #1

  1. JevHosted API, generally availableWhere it runsdecider-4bSelf-hosted on one GPU, about 8.4 GB in bf16
  2. JevProprietary, hostedLicencedecider-4bApache-2.0 weights and package
  3. Jev83.3 · 0.65s median over the networkSpeeddecider-4b92.9 · 17ms raw on the benchmark’s GPU
  4. Jev94.5% correctEvaluation-style casesdecider-4b87.7% correct
  5. Jev74.1% correctHardest test casesdecider-4b67.3% correct
  6. Jev36.7% correctSealed decisionsdecider-4b34.7% correct
  7. Jev64k tokens per requestContextdecider-4b32k tokens

Which decider?

Mapika publishes several decider models, and releases come quickly. JevBench has ranked three of them; this page compares Jev with the one in first place.

  1. decider-4b v2by Mapika

    Qwen3.5-4B-Base, 4.2B parameters, about 8.4 GB in bf16. Released on 24 September; the version JevBench measured, kept under the Hugging Face tag v2.

    #1 of 89 at 64.1

  2. decider-4b v2.1by Mapika

    The current default, released the same day: v2 plus a further LoRA stage. Its author reports it slightly weaker on JevBench’s public hard items, 0.649 against 0.676, and less well calibrated on hard items.

    Not ranked in JevBench v1.4.2

  3. decider-35b-a3bby Mapika

    Qwen3.5-35B-A3B-Base, 34.7B parameters with 3B active: about 65 GB in bf16 or 19.6 GB in NVFP4.

    #19 of 89 at 41.2, held back by cost

  4. decider-2bby Mapika

    Qwen3.5-2B-Base, 1.9B parameters, about 4 GB. The first decider model, carried into v1.4.2 from an earlier measurement.

    #39 of 89 at 30.7, held back by calibration

  5. decider-0.8b and decider-2b-visionby Mapika

    A smaller text model and a vision-language variant that reads decisions from pixels, still on older text weights.

    Not in JevBench v1.4.2

JevBench v1.4.2, measured the same way

One benchmark measured all 93 systems under one method and ranked 89, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.

Composite score

#1 decider-4b64.1
#2 Jev63.3
#3 JevK562.0
#4 Cygnet61.8
#5 Hopper59.4
#6 Winnow55.6
#8 djev52.2
#11 SemIf47.7
#27 OpenJev36.9
#41 Laya30.3

The top four finish within 2.3 points of each other, and then the board falls away sharply. The composite hides where systems actually differ, so the two charts below break it apart.

Capability, higher is better

IntelligenceHow often it picks the right answer

Jev53.1
decider-4b49.4

CalibrationWhether 0.8 really means about 80%

Jev76.3
decider-4b75.0

SpeedMeasured response time

Jev83.3
decider-4b92.9

Accuracy by how hard the decision is

Easy72 straightforward cases

Jev100%
decider-4b100%

Standard96 everyday cases

Jev99%
decider-4b96.9%

Judge146 evaluation-style calls

Jev94.5%
decider-4b87.7%

Hard220 genuinely ambiguous cases

Jev74.1%
decider-4b67.3%

Sealed308 private cases, new in v1.4; chance is 29.3%

Jev36.7%
decider-4b34.7%

Easy and standard decisions separate almost nothing. The hard tier, highlighted, is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for.

One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.

Every system on the board

The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and decider-4b are highlighted.

ModelRankScoreMeasured latencyRuns on
decider-4b v2 (Mapika, Qwen3.5-4B)#164.10.02s raw, 0.18s adjustedSelf-hosted, RTX PRO 6000
Jev 1.13.0 (TypeSafe)#263.30.65s median / 0.72s p95Hosted production API
JevK5 v0.2.0 (allebee, Qwen3.5-4B)#362.0Not re-measured in v1.4Self-hosted, RunPod GPU
Cygnet (blockbrain, frozen Gemma 4 12B)#461.80.04s raw, 0.22s adjustedSelf-hosted, RTX PRO 6000
Hopper (HopitAI, Qwen3.5-4B LoRA)#559.40.13s raw, 0.41s adjustedSelf-hosted, RTX A6000
Winnow-12B Q8 (EldanRing, Gemma 4 12B)#655.60.23s raw, 0.60s adjustedSelf-hosted, RTX 4090
djev (Maisa, DiffusionGemma)#852.20.24s median / 0.31s p95Hosted API, preview
SemIf (Qwen3.5-4B)#1147.70.20s raw, 0.55s adjustedSelf-hosted, RunPod GPU
OpenJev (razorback16)#2736.90.24s raw, 0.63s adjustedSelf-hosted, RunPod GPU
Laya (ModernBERT-large)#4130.30.79s raw, 1.72s adjustedSelf-hosted, CPU

JevBench v1.4.2, published on 25 September 2026, lists 93 systems and ranks 89 of them. Each is scored on the same 534 public decisions (72 easy, 96 standard, 146 judge, 220 hard) plus 308 sealed decisions whose text stays private, one request at a time. The composite is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard. Sealed decisions make up 20% of intelligence, and a system that does more than 25 points better on public items than on sealed ones loses intelligence in proportion. Every Jev-style system scores far lower on the sealed set than on the public one, Jev included (36.7% against a 29.3% chance level). Version 1.4.2 kept the v1.4.0 formulas and every earlier measurement; it added 17 systems over two releases, which is why ranks moved while scores did not. Self-hosted and demo latency is adjusted by x2 plus 0.15s to approximate production load, so those numbers are not comparable with the raw latency their authors publish.

Jev and decider-4b feature by feature

AttributeJev 1.13.0decider-4b
What it isHosted System One model from TypeSafe, version jev-1.13.0Open-weight Qwen3.5-4B-Base fine-tuned for one-pass typed decisions
LicenceProprietary, hostedApache-2.0 for weights and package
Request formatPOST /v1/systemoneThe same wire format; TypeSafe’s SDKs work with a changed base URL
Question typesNoul, choice, score — many per call, answered in parallelNoul, choice up to 255 options, score over 2–10 levels — one answer slot per question in one pass
Context64k tokens per request32k tokens
LanguagesBest in English; other languages work with lower accuracyEnglish only
CalibrationPost-trained with RLCD; JevBench 76.3One fitted temperature, 1.935 in the measured build; JevBench 75.0
HardwareNone — it is an API callOne CUDA GPU, about 8.4 GB in bf16; also runs on CPU
Version measuredjev-1.13.0v2, kept under the Hugging Face tag v2; v2.1 is now the default

What each one is better at

Where Jev wins

  • Intelligence 53.1 against 49.4, the highest in the top ten; first on JevBench’s intelligence-only view, where decider-4b is fifth.
  • Hard-tier accuracy 74.1% against 67.3%, judge tier 94.5% against 87.7%, sealed set 36.7% against 34.7%.
  • Paraphrase robustness on the sealed set, 64.3% against 50.0%, and clear leads on probability and multi-hop questions.
  • Twice the context, 64k tokens against 32k, and usable beyond English.
  • Nothing to operate, and one stable hosted version rather than a model with three releases between 22 and 24 September.

Where decider-4b wins

  • First of 89 in JevBench v1.4.2 at 64.1 against Jev’s 63.3, and first again when the benchmark weights speed or cost more heavily.
  • Speed 92.9 against 83.3: a 17ms raw median per decision on the benchmark’s RTX PRO 6000.
  • Apache-2.0 weights, with the training code and the builders for its public datasets in the repository.
  • Ahead of Jev on the sealed set’s ambiguous-abstain family, 43.2% against 29.7%.
  • Speaks TypeSafe’s wire format, so moving a Jev integration is a base-URL change.

The analysis

Where Jev and decider-4b really differ, and why the bars above land where they do.

decider-4b v2 is the first system to finish above Jev on JevBench. Version 1.4.2 scores it 64.1 against Jev’s 63.3, first of 89 ranked systems, and the benchmark’s independent check before listing it first found the result legitimate. The lead comes from two axes: speed, 92.9 against 83.3, with a raw median of 17ms per decision on the benchmark’s own GPU, and the benchmark’s cost estimate, where a 4B model is cheap to run.

On the axes that measure answers, Jev is ahead. Intelligence is 53.1 against 49.4 and calibration 76.3 against 75.0. Jev gets 74.1% of the 220 hard decisions right against 67.3%, 94.5% of the judge tier against 87.7%, and 36.7% of the 308 sealed decisions against 34.7%. Weigh the axes differently and the order changes: JevBench’s own intelligence-only view, which keeps its speed and cost gates, puts Jev first and decider-4b fifth, and with accuracy weighted at 60% Jev is first and decider-4b second.

So the composite and the capability ranking disagree, and both are honest. If your decisions are high-volume and mostly routine, and you can run a GPU, decider-4b gives you most of Jev’s accuracy in milliseconds on your own hardware. If the cases that matter are the hard ones, Jev is still the more accurate model.

What decider-4b actually is

decider-4b v2, the open System One rebuild from Mapika, by Mapika.

decider is a family of open models from Mapika that reproduce the System One model class without generating text. The 4B model is Qwen3.5-4B-Base fine-tuned to read a state and a set of typed questions (choice over 2 to 255 options, score over 2 to 10 levels, and noul, the probability of yes) and to return one probability distribution per question from a single forward pass. It is Apache-2.0 and installs from PyPI as decider-ai. Its README states that it is not affiliated with TypeSafe and that nothing was distilled from Jev: the training mixture is public datasets plus data labelled by a local Qwen3.5-27B teacher.

Its server speaks TypeSafe’s /v1/systemone wire format, and the README says TypeSafe’s own SDKs work unchanged when pointed at it. The 4B model needs one CUDA GPU with about 8.4 GB in bf16, and the server captures its CUDA graphs at start-up. The package also runs on CPU, and Apple Silicon support covers the smaller dense models.

The version JevBench ranked is v2, released on 24 September: v1 plus a LoRA stage on harder decisions. The same day Mapika released v2.1, which recovers some game-playing ability but, by the author’s own measurement, is slightly weaker on JevBench’s public hard items and less well calibrated on them. If you want the model that holds first place, pin the Hugging Face tag v2.

What the benchmark disclosed about its training

JevBench records one disclosure from the author: 8,000 of the rows in the v2 training stage came from generators written from the published names of the ten sealed-question families, without reading any sealed item. The benchmark’s independent review judged the result legitimate, and notes that the author’s private second-stage training rows could not be audited for overlap with the public items.

Neither point means the score is wrong. They do mean decider-4b v2 was built after the sealed families were named, which is worth knowing when you read its sealed-set numbers.

Where each one wins on the sealed set

Grouped by family, decider-4b is ahead on ambiguous questions where abstaining is the right call, 43.2% against 29.7%, and slightly ahead on long policies, 30.0% against 27.5%, and on dates and numbers, 30.4% against 28.6%. The two tie on safety judgments at 37.5%.

Jev leads on paraphrase robustness, 64.3% against 50.0%; probability questions, 50.0% against 35.7%; multi-hop lookups, 44.7% against 34.2%; trap and adversarial items, 41.7% against 33.3%; trade-offs, 38.5% against 34.6%; and hard judging, 34.1% against 31.7%. Every system of this kind scores low on the sealed set overall, not far above the 29.3% chance level, so treat these as a guide to what to test rather than a verdict.

So which should you use?

Pick Jev if

  • The decisions that matter are the hard ones: long policies with exceptions, multi-hop lookups, questions reworded by different people.
  • Inputs run past 32k tokens, or are not in English.
  • You want a hosted, versioned model with published rate limits and nothing to serve.

Pick decider-4b if

  • You already run GPUs and want millisecond decisions on your own hardware.
  • Most of your decisions are routine, where the two are close, and volume makes per-decision cost matter.
  • You need open weights for offline use, data residency or fine-tuning.

Try Jev on the cases you are actually arguing about

decider-4b is close enough, and fast enough, that the honest test is your own cases. Run them on Jev here with no setup and keep the answers: if decider-4b agrees where it matters, you have a strong case for self-hosting, and if it does not, you will see exactly which cases it misses.

Jev vs decider-4b FAQ

Is decider-4b better than Jev?

On JevBench’s composite, yes, by 0.8 points: 64.1 against 63.3. The composite weighs speed and cost equally with intelligence and calibration. On accuracy Jev is ahead: intelligence 53.1 against 49.4, hard decisions 74.1% against 67.3%, sealed decisions 36.7% against 34.7%.

Is decider made by TypeSafe?

No. It is an independent Apache-2.0 project by Mapika. Its README states that it is not affiliated with or endorsed by TypeSafe AI and that nothing was distilled from Jev.

Can I point the TypeSafe SDK at decider?

Yes, according to its README: the server implements POST /v1/systemone in TypeSafe’s wire format, so the SDKs work once their base URL points at your server. Check your input lengths against its 32k-token context first.

Which decider version is ranked first?

decider-4b v2, released on 24 September and still available under the Hugging Face tag v2. The current default, v2.1, was released the same day and is not ranked by JevBench; its author reports it slightly weaker on JevBench’s public hard items.

What about decider-2b and decider-35b-a3b?

Both are ranked in JevBench v1.4.2: decider-35b-a3b at #19 with 41.2, held back by the cost of a 35B model, and decider-2b at #39 with 30.7, held back by calibration.

Other Jev comparisons

Every number on these pages is quoted from a published source and was read on September 26, 2026.

Sources

decider-4b and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on September 26, 2026 and can change.