Jev AI

Model guide · Open System One rebuild · Qwen3.5-4B-Base · pip install decider-ai · updated October 7, 2026

decider-4b

Millisecond decisions from 8.4 GB of open weights, on TypeSafe’s wire format.

decider-4b reads a state and a set of typed questions and returns one probability distribution per question from a single forward pass of an open Qwen3.5-4B model. Its server speaks TypeSafe’s /v1/systemone wire format, so TypeSafe’s own SDKs work when pointed at it. It is Apache-2.0, installs from PyPI, and was the first system to finish above Jev on JevBench when it was added; on the re-measured v1.6.0 pool it is #17.

decider-4b is an independent project by Mapika. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.

Developer
Mapika (independent)
Base model
Qwen3.5-4B-Base, 4.2B parameters
Weights
About 8.4 GB in bf16, Mapika/decider-4b
Licence
Apache-2.0 weights and package
Install
pip install decider-ai
Context
32k tokens
Questions
Choice over 2–255 options · score over 2–10 levels · noul
Runs on
One CUDA GPU; also CPU; Apple silicon for smaller sizes

decider-4b on JevBench v1.6.0

One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.

decider-4b

41.2rank #17 · Capability #9

Self-hosted, RTX 5090 · 0.02s raw, 0.20s adjusted

Jev 1.13.0

71.1rank #2 · Capability #3

Hosted API · 0.24s median / 0.30s p95

IntelligenceHow often it picks the right answer, half from sealed decisions

decider-4b40.1
Jev62.4

CalibrationWhether 0.8 really means about 80%

decider-4b89.6
Jev90.6

SpeedMeasured response time

decider-4b91.8
Jev91.5

Open half 45.2 · sealed half 35.0; chance-corrected by request type: choice 65.0, yes/no 19.1, score 36.3. The composite also weighs a fourth axis, cost, which these bars leave out. Self-hosted latency is adjusted by the benchmark to approximate production load.

What decider-4b is

decider is a family of open models from Mapika that reproduce the System One model class without generating text. The 4B model is Qwen3.5-4B-Base fine-tuned to read a state and typed questions and return one probability distribution per question in one pass. Its README states that it is not affiliated with TypeSafe and that nothing was distilled from Jev: the training mixture is public datasets plus data labelled by a local Qwen3.5-27B teacher.

The server implements TypeSafe’s /v1/systemone wire format and the README says TypeSafe’s SDKs work unchanged with a changed base URL. The 4B needs one CUDA GPU with about 8.4 GB in bf16 and captures its CUDA graphs at start-up; the package also runs on CPU, and Apple Silicon support covers the smaller dense models.

JevBench records one disclosure from the author: 8,000 rows in the v2 training stage came from generators written from the published names of the ten sealed-question families, without reading any sealed item. The benchmark’s independent review judged the result legitimate.

Where it stands on JevBench

JevBench v1.6.0 ranks decider-4b v2 #17 of 92 at 41.2, measured on an RTX 5090 at a 25ms raw median per decision. Speed is 91.8, level with Jev’s 91.5, and calibration 89.6 against Jev’s 90.6. Intelligence is 40.1 against Jev’s 62.4, and 35.0 on the sealed half against 63.6; its open-minus-sealed gap of 10.2 points is about four times the field median. Because intelligence sits below the benchmark’s floor of 50, the composite is scaled down, which is why the score is 41.2 while the individual axes look respectable. On the Capability ranking, which leaves speed and cost out, it is ninth.

Which decider?

Mapika publishes several decider models and releases come quickly. JevBench v1.6.0 ranks three of them.

  1. decider-4b v2by Mapika

    Qwen3.5-4B-Base, 4.2B parameters, about 8.4 GB in bf16. Released on 24 September; the version JevBench measured, kept under the Hugging Face tag v2.

    #17 of 92 at 41.2

  2. decider-4b v2.1by Mapika

    The current default, released the same day: v2 plus a further LoRA stage. Its author reports it slightly weaker on JevBench’s public hard items, 0.649 against 0.676, and less well calibrated on hard items.

    Not measured in JevBench v1.6.0

  3. decider-35b-a3bby Mapika

    Qwen3.5-35B-A3B-Base, 34.7B parameters with 3B active: about 65 GB in bf16 or 19.6 GB in NVFP4.

    #29 of 92 at 24.8, held back by cost

  4. decider-2bby Mapika

    Qwen3.5-2B-Base, 1.9B parameters, about 4 GB. The first decider model.

    #55 of 92 at 7.7, held back by intelligence

  5. decider-0.8b and decider-2b-visionby Mapika

    A smaller text model and a vision-language variant that reads decisions from pixels, still on older text weights.

    Not measured in JevBench v1.6.0

Run decider-4b

Install the package from PyPI, start the server on a CUDA GPU, and point a TypeSafe SDK or plain HTTP client at it. Pin the Hugging Face tag v2 if you want the exact build JevBench ranked.

pip install decider-ai
# Start the server per the README; the 4B weights download from huggingface.co/Mapika/decider-4b
# (tag v2 is the build JevBench measured; v2.1 is the current default).
curl http://localhost:8000/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"state": "Review: The app crashes every time I open the camera tab.",
       "questions": {"sentiment": {"type": "score", "instructions": "How negative is this review?",
                                   "criteria": ["neutral", "mildly negative", "very negative"]}}}'
  • The server captures CUDA graphs at start-up, so the first request after launch is slower than the 25ms raw median JevBench measured.
  • English only; inputs are capped at 32k tokens.
  • Three releases shipped between 22 and 24 September 2026; pin a tag in production.

Limits to plan around

From the project’s own documentation and JevBench v1.6.0.

  • Intelligence 40.1 in JevBench v1.6.0, with a 10.2-point gap between open and sealed decisions.
  • English only, 32k-token context.
  • A 4B model with frequent releases: v2 and v2.1 shipped the same day and score differently.

Measure it against Jev on your own cases

decider-4b is close enough, and fast enough, that the honest test is your own cases. Run them on Jev here with no setup and keep the answers: if decider-4b agrees where it matters, you have a strong case for self-hosting, and if it does not, you will see exactly which cases it misses.

decider-4b FAQ

What is decider-4b?

An Apache-2.0 open model from Mapika that reproduces the System One model class on Qwen3.5-4B-Base: it reads a state and typed questions and returns a probability distribution per question in one forward pass, behind TypeSafe’s /v1/systemone wire format. It installs from PyPI as decider-ai.

Is decider made by TypeSafe?

No. Its README states that it is not affiliated with or endorsed by TypeSafe AI and that nothing was distilled from Jev.

Which decider version does JevBench rank?

decider-4b v2, released on 24 September 2026 and available under the Hugging Face tag v2. The current default, v2.1, was not measured on the v1.6.0 pool; its author reports it slightly weaker on JevBench’s public hard items.

Can I point the TypeSafe SDK at decider?

Yes, according to its README: the server implements POST /v1/systemone in TypeSafe’s wire format, so the SDKs work once their base URL points at your server. Check input lengths against the 32k-token context first.

What hardware does decider-4b need?

One CUDA GPU with about 8.4 GB free in bf16. The package also runs on CPU, and Apple Silicon support covers the smaller dense models. JevBench measured it on an RTX 5090.

More model guides

Every figure on this page is quoted from a published source and was read on October 7, 2026.

Sources

decider-4b and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.