Model guide · Jev-compatible decision server · DiffusionGemma 26B-A4B · reads images · updated October 7, 2026
OpenJev
One name, ten projects. Find yours, then run the drop-in.
openJev is not a model. It is a name that at least ten unrelated projects have taken, by ten different authors, and JevBench v1.6.0 ranks seven of them anywhere from #40 to #92. This guide lists every one so you can find the project you actually meant, then documents the one built to replace Jev directly: razorback16’s server on DiffusionGemma, which answers on the /v1/systemone wire format, accepts the model name jev-latest and reads images.
OpenJev is an independent project by razorback16 / Codiv. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.
- Developer
- razorback16; hosted by Codiv
- Model
- DiffusionGemma 26B-A4B-it, NVFP4, about 18 GB
- Licence
- Apache-2.0; weights from NVIDIA and Google
- Wire format
- /v1/systemone; accepts openjev-latest, jev-latest, jev-preview
- Inputs
- Text, plus up to 8 images per request
- Questions
- Noul · choice up to 128 options · score over 2–10 levels
- Hardware
- 24 GB NVIDIA GPU under vLLM, or Apple silicon with 16 GB free under MLX
- Hosted
- api.codiv.ai/v1/systemone
OpenJev on JevBench v1.6.0
One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.
JevBench v1.6.0 measured razorback16’s OpenJev only in thinking mode, where it ranks #40 at 17.3 with intelligence 70.3 and calibration 81.4. Its default configuration was not measured, so there is no like-for-like bar to draw against Jev; the standings of every other openJev are in the list below.
What razorback16’s OpenJev is
It runs DiffusionGemma 26B-A4B-it quantised to NVFP4 under an Apache-2.0 licence. A diffusion model denoises many tokens at once rather than writing them one after another, and OpenJev turns that into a decision engine: it builds a masked canvas in which only the answer slots are uncertain, takes one read-only forward pass and reads the probability distribution over the label tokens. The model never writes into those slots, so an answer cannot go off-schema. Above 0.1 entropy it re-reads up to four times with fresh noise and averages.
It supports the same three primitives as Jev, and adds up to eight images per request, a tunable 1–8 denoising steps, 1–32 samples per read and a thinking budget of up to 4,096 tokens. Images cannot be combined with thinking mode, and the Apple silicon backend does not return reasoning content separately.
Confidence is computed as 1 − H(p)/ln K: one when the model is certain, zero when the distribution is flat. That measures how concentrated the answer is, not how often it is right, and the documentation publishes no independent calibration check. Validate thresholds on your own labelled data before letting it auto-approve anything.
Where it stands on JevBench
JevBench v1.6.0 measured razorback16’s OpenJev only in thinking mode. There it beats Jev on intelligence, 70.3 against 62.4, with calibration 81.4 against Jev’s 90.6, but ranks #40 of 92 at 17.3 because thinking needs far more compute per decision: its cost score is 28.5 and its speed 72.6, and the harmonic mean punishes weak axes hard. The default configuration was not measured; the last release that did put it well below Jev on intelligence and calibration.
Which openJev do you mean?
Every project below has been called openJev or Open-Jev by its author or by someone writing about it. They share a name and an idea, and nothing else: different base models, sizes, licences and quality.
OpenJevby razorback16 / Codiv
Jev-compatible decision server on DiffusionGemma 26B-A4B (NVFP4, about 18 GB). Accepts jev-latest, so TypeSafe SDKs work unchanged. Reads images. Hosted by Codiv, or self-host on a 24 GB GPU or Apple silicon. This guide.
Default not measured · #40 of 92 at 17.3 in thinking mode
openjev-sglangby ekzhang
Qwen3.6-35B-A3B served on SGLang, also implementing the /v1/systemone wire format. Its documentation states that its probabilities are not calibrated estimates of correctness.
Not measured in JevBench v1.6.0
Open-Jev 9B and 2Bby Zefan Cai
LoRA adapters plus a trained scalar decision head and calibration temperature over a Qwen backbone. Not merged checkpoints: they need the pinned upstream weights and the Open-Jev loader.
#41 at 17.0 and #61 at 3.2
openJev Verdictby heman10x
A 151M ModernBERT-base encoder with a GLiClass head, post-trained with RLCD and temperature-calibrated. Near-perfect on the data it was fitted to, weak on anything else. Runs on CPU or in a browser tab.
#74 at 0.2 (version 1.4) and #79 at 0.1
open-jev-deberta-v3-largeby Kotoba Labs
A DeBERTa-v3 encoder run as a local CPU decision model.
#82 of 92 at 0.0
SemIfby TheoLeeCJ
Was called OpenJev until it was renamed. Reads typed option logits out of a frozen Qwen3.5-4B. Older benchmark tables and blog posts still use the old name.
#45 of 92 at 15.5 — see the SemIf guide
OpenSourceJevby sabeel111
A Qwen3.5-4B rebuild quantised to Q4_K_M and served with llama.cpp. MIT code over Apache-2.0 weights.
#52 of 92 at 10.3
Open Jev JSON Canvasby JoshuaSP
DiffusionGemma 26B-A4B filling a JSON answer canvas in one denoising step. It returns a label rather than a probability distribution, so calibration counts as zero.
#92 of 92 at 0.0; intelligence 61.2
OpenJevby zhangcy122
A separate typed decision API over open LLMs using constrained logprob calibration. Same name as razorback16’s, unrelated codebase.
Not in JevBench v1.6.0
open-jev-typed-decision-engineby eightman999
A 150M typed decision engine that trains on a Colab T4 in about 30 minutes. The author reports 0.697 against Jev’s 0.727 on their own set.
Not in JevBench v1.6.0
open-alternative-jevby IkerMoel
Not an openJev, but constantly mistaken for one: a Qwen3.5-4B rebuild with a similar name and a similar goal.
#53 of 92 at 9.4
Open Spark Jev (spark-s1)by abhishek085
Not an openJev either, though the name is close: local Qwen3.5-4B decision models built for NVIDIA DGX Spark, Apache-2.0. The measured model is spark-s1-4b-v6.
#11 of 92 at 44.9
Run OpenJev
Self-host under vLLM on a 24 GB NVIDIA GPU or under MLX on Apple silicon, or call Codiv’s hosted endpoint. Because it accepts jev-latest on the /v1/systemone wire format, a TypeSafe SDK only needs its base URL changed.
git clone https://github.com/razorback16/openjev
cd openjev
# Follow README.md: vLLM backend (24 GB NVIDIA GPU) or MLX backend (Apple silicon, 16 GB free).
# The NVFP4 weights are about 18 GB. Hosted alternative: https://api.codiv.ai/v1/systemone
curl http://localhost:8000/v1/systemone \
-H "Content-Type: application/json" \
-d '{"model": "jev-latest",
"state": "Listing photo attached. Description: brand new, sealed box.",
"questions": {"matches": {"type": "noul", "instructions": "Does the description match the photo?"}}}'- vLLM on a 24 GB GPU handles up to 64 in-flight reads at about 94ms p50 at concurrency 1; MLX on an M3 Ultra runs 0.2–0.4s per read, one at a time.
- Accuracy, cost and latency all move with denoising steps, sample count and thinking budget; fix them before you benchmark.
- Confidence is derived from entropy, so validate thresholds on labelled data.
Limits to plan around
From the project’s own documentation and JevBench v1.6.0.
- Up to 128 options per choice, against Jev’s 255.
- Images cannot be combined with thinking mode.
- No independent calibration check published; the thinking build scores 81.4 in JevBench v1.6.0.
- An 18 GB quantised checkpoint and a 24 GB GPU, or an Apple silicon machine, to keep serving.
Measure it against Jev on your own cases
Whichever openJev you landed on, the only comparison that settles it is on your own awkward cases. Running Jev here takes a minute with no setup, and the answers come back with the full probability distribution — which is exactly the thing most of the projects above tell you not to trust yet.
OpenJev FAQ
Which project is the real openJev?
None of them is official, and there are at least ten. JevBench v1.6.0 ranks seven: razorback16’s DiffusionGemma server, in thinking mode only, is highest at #40; Zefan Cai’s Open-Jev 9B and 2B are #41 and #61; SemIf, which used to carry the name, is #45; OpenSourceJev is #52; heman10x’s Verdict builds are #74 and #79; Kotoba Labs’ DeBERTa build is #82; and JoshuaSP’s JSON canvas is #92.
Can I point the TypeSafe SDK at OpenJev?
At razorback16’s, yes: it implements the /v1/systemone wire format and accepts jev-latest and jev-preview as model names, so changing the base URL is enough. openjev-sglang does the same. The others have their own interfaces.
Is OpenJev as accurate as Jev?
Its default configuration was not measured in JevBench v1.6.0. In thinking mode it is more accurate than Jev on intelligence, 70.3 against 62.4, but needs far more compute per decision and ranks #40 at 17.3.
What hardware does OpenJev need?
About 18 GB of NVFP4 weights: a 24 GB NVIDIA GPU under vLLM, or Apple silicon with 16 GB free under MLX. Codiv also hosts it at api.codiv.ai.
Can OpenJev read images?
Yes, up to eight per request, though not together with thinking mode. Jev is text only.
Is openJev Verdict a good Jev replacement?
Only on data it was fitted to. On its own held-out set it reports 95% top-1 accuracy, but its author measures 48.07% on TypeSafe’s 337-case benchmark against 90.80% for Jev, and in JevBench v1.6.0 both Verdict builds answer the sealed decisions barely above chance.
More model guides
Every figure on this page is quoted from a published source and was read on October 7, 2026.
- SemIfSemantic ifs from a frozen open model, with no generation at all.
- djevDecisions read off a diffusion model in one denoising step.
- Imajev-4BJev’s request format, open weights, and two photos per request.
- Jev 1.13.0The hosted model every project here is measured against.
- Jev-OmniGemma 4 12B decision model with image input.
Sources
OpenJev and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.