Jev AI

Model guide · Open-weight Jev alternative · Qwen3.5-4B + distilled LoRA · updated October 7, 2026

JevK5

The fastest open 4B on the board, with the weights published.

JevK5 answers the same yes/no, choice and score questions as Jev from open Qwen3.5-4B weights with a distilled LoRA merged in. It reads every option’s probability from the next-token logits in one forward pass on your own GPU, accepts the TypeSafe-style /v1/systemone request and is published under Apache-2.0. Version 0.3 is the current release and the one JevBench v1.6.0 ranks.

JevK5 is an independent project by allebee. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.

Developer
allebee (independent)
Base model
Qwen3.5-4B, LoRA merged into the weights
Weights
alibiserikbay/JevK5 on Hugging Face
Licence
Apache-2.0 weights and code
Inputs
Text, English only
Context
Inputs over 16,384 tokens are refused
Options per choice
Up to 16
GPU memory
About 9 GB in bf16

JevK5 on JevBench v1.6.0

One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.

JevK5

37.4rank #18 · Capability #10

Self-hosted, H100 · 0.02s raw, 0.19s adjusted

Jev 1.13.0

71.1rank #2 · Capability #3

Hosted API · 0.24s median / 0.30s p95

IntelligenceHow often it picks the right answer, half from sealed decisions

JevK538.5
Jev62.4

CalibrationWhether 0.8 really means about 80%

JevK589.9
Jev90.6

SpeedMeasured response time

JevK593.9
Jev91.5

Open half 38.3 · sealed half 38.7; chance-corrected by request type: choice 59.7, yes/no 21.7, score 34.2. The composite also weighs a fourth axis, cost, which these bars leave out. Self-hosted latency is adjusted by the benchmark to approximate production load.

What JevK5 is

JevK5 reads answers the way SemIf does: a softmax over the next-token logits of the answer letters, with one fitted temperature, so every option gets a probability in a single pass and nothing is generated. It states that it is not affiliated with TypeSafe and does not reproduce Jev’s unpublished architecture.

The training is what separates it from a plain logit reader. Qwen3.6-27B, with thinking, wrote documents carrying hard typed questions across 17 business domains and 11 decision families, answered each question twice, and kept it only when both answers matched. JevK5 was trained on 3,272 of those questions plus 3,272 human-labelled items from public datasets such as MMLU-Pro, BoolQ and banking77. The author states that no JevBench item and no Jev output was used.

The author documents a correction made on 23 September: a small hand-written set used to choose the calibration setting echoed the wording of a few public benchmark items. It was rewritten, and the author reports that no accuracy figure changed.

Versions and where they stand

The repository appeared on 22 September 2026 and version 0.2.0 followed. JevBench v1.6.0 ranks version 0.3 #18 of 92 at 37.4, with intelligence 38.5, calibration 89.9 and the highest speed score of any system charted on this site, 93.9. Version 0.2 is #23 at 32.1. Jev is #2 at 71.1.

Plumb-4B, a LoRA fine-tune of JevK5 v0.2 by another author, ranks #10 at 48.3, eight places above JevK5 itself: the base takes further fine-tuning well. On Benchmark Heaven’s Capability ranking, which averages intelligence and calibration, JevK5 v0.3 is tenth.

Run JevK5

JevK5 installs from GitHub as a Python package and ships a server that accepts the TypeSafe-style request shape. In bf16 it needs about 9 GB of GPU memory.

git clone https://github.com/allebee/jevk5
cd jevk5
# Follow README.md to install the package and start the server.
# The weights download from huggingface.co/alibiserikbay/JevK5 on first run.
curl http://localhost:8000/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"state": "Ticket: I was charged twice for the same month.",
       "questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
                              "criteria": {"billing": null, "technical": null, "account": null}}}}'
  • On one H100 the author measures a 13.5ms median on easy and standard items and a 30ms median on hard items with 1–4k-token documents.
  • Without the optional flash-linear-attention package, long inputs run slower.
  • Requests are serialized on one GPU; questions in one call are evaluated separately rather than in parallel.

Limits to plan around

From the project’s own documentation and JevBench v1.6.0.

  • English only.
  • Up to 16 options per choice question; Jev takes up to 255.
  • Inputs over 16,384 tokens are refused; Jev accepts 64k.
  • The author says its quality on Jev’s published real-world workflows has not been measured yet.

Measure it against Jev on your own cases

JevK5 is close enough that the honest comparison is on your own cases. Run them on Jev here — it takes a minute with no setup — and keep the answers: if JevK5 agrees on the cases you care about, you have a strong reason to self-host, and if it does not, you know exactly where the gap is.

JevK5 FAQ

What is JevK5?

An open-weight Jev alternative by allebee: Qwen3.5-4B with a distilled LoRA merged into the weights, published on Hugging Face as alibiserikbay/JevK5 under Apache-2.0. It answers noul, choice and score questions in one forward pass and ships a server that accepts the TypeSafe-style request.

Is JevK5 made by TypeSafe?

No. It is an independent project and states that it is not affiliated with TypeSafe AI. It does not reproduce Jev’s architecture.

Which JevK5 version should I run?

Version 0.3, the current release. JevBench v1.6.0 ranks it #18 at 37.4 against version 0.2 at #23, with better intelligence, 38.5 against 36.3, and better calibration.

What hardware does JevK5 need?

About 9 GB of GPU memory in bf16. The author’s published figures are on an H100 and an L40S.

Can I send Jev requests to JevK5?

Its server accepts the TypeSafe-style /v1/systemone request shape, so requests written for Jev need little or no change. Check the 16-option and 16,384-token limits first; the project does not claim full SDK compatibility.

How is JevK5 different from SemIf and Plumb-4B?

It uses SemIf’s option-logit readout on the same Qwen3.5-4B base, then adds a distilled LoRA and a fitted temperature. Plumb-4B is a further LoRA fine-tune of JevK5 v0.2 by another author and ranks #10, above JevK5’s #18; SemIf is #45.

More model guides

Every figure on this page is quoted from a published source and was read on October 7, 2026.

Sources

JevK5 and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.