On this site
Jev 1.13.0
TypeSafe · System One model
Fourth overall, and still the highest intelligence score in the top ten.
63.3JevBench v1.4.2.2 rank #4
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026
The new first place on JevBench, from a 4B model that also reads photos.
On this site
TypeSafe · System One model
Fourth overall, and still the highest intelligence score in the top ten.
63.3JevBench v1.4.2.2 rank #4
vs
Alternative
Mohit Garg · Open Qwen3.5-4B decision model that also reads photos
First of 91, with the best calibration in the top ten.
67.4JevBench v1.4.2.2 rank #1
One benchmark measured all 95 systems under one method and ranked 91, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.
The top four finish within 4.1 points of each other, and then the board falls away sharply. The composite hides where systems actually differ, so the two charts below break it apart.
IntelligenceHow often it picks the right answer
CalibrationWhether 0.8 really means about 80%
SpeedMeasured response time
Easy72 straightforward cases
Standard96 everyday cases
Judge146 evaluation-style calls
Hard220 genuinely ambiguous cases
Sealed308 private cases, new in v1.4; chance is 29.3%
Easy and standard decisions separate almost nothing. The hard tier, highlighted, is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure for that tier.
One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.
The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and Imajev-4B are highlighted.
| Model | Rank | Score | Measured latency | Runs on |
|---|---|---|---|---|
| Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA) | #1 | 67.4 | 0.04s raw, 0.23s adjusted | Self-hosted GPU |
| Plumb-4B (crh225, JevK5 v0.2 + LoRA) | #2 | 65.8 | See the source table | Self-hosted GPU |
| decider-4b v2 (Mapika, Qwen3.5-4B) | #3 | 64.1 | 0.02s raw, 0.18s adjusted | Self-hosted, RTX PRO 6000 |
| Jev 1.13.0 (TypeSafe) | #4 | 63.3 | 0.65s median / 0.72s p95 | Hosted production API |
| JevK5 v0.2.0 (allebee, Qwen3.5-4B) | #5 | 62.0 | Not re-measured in v1.4 | Self-hosted, RunPod GPU |
| Cygnet (blockbrain, frozen Gemma 4 12B) | #6 | 61.8 | 0.04s raw, 0.22s adjusted | Self-hosted, RTX PRO 6000 |
| Hopper (HopitAI, Qwen3.5-4B LoRA) | #7 | 59.4 | 0.13s raw, 0.41s adjusted | Self-hosted, RTX A6000 |
| Winnow-12B Q8 (EldanRing, Gemma 4 12B) | #8 | 55.6 | 0.23s raw, 0.60s adjusted | Self-hosted, RTX 4090 |
| djev (Maisa, DiffusionGemma) | #10 | 52.2 | 0.24s median / 0.31s p95 | Hosted API, preview |
| SemIf (Qwen3.5-4B) | #13 | 47.7 | 0.20s raw, 0.55s adjusted | Self-hosted, RunPod GPU |
| OpenJev (razorback16) | #29 | 36.9 | 0.24s raw, 0.63s adjusted | Self-hosted, RunPod GPU |
| Laya (ModernBERT-large) | #43 | 30.3 | 0.79s raw, 1.72s adjusted | Self-hosted, CPU |
JevBench v1.4.2.2, published on 27 September 2026, lists 95 systems and ranks 91 of them. Each is scored on the same 534 public decisions (72 easy, 96 standard, 146 judge, 220 hard) plus 308 sealed decisions whose text stays private, one request at a time. The composite is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard. Sealed decisions make up 20% of intelligence, and a system that does more than 25 points better on public items than on sealed ones loses intelligence in proportion. Every Jev-style system scores far lower on the sealed set than on the public one, Jev included (36.7% against a 29.3% chance level). The scoring formulas have not changed since v1.4.0 and no earlier measurement has been redone; 19 systems were added over four releases, which is why ranks moved while scores did not. Self-hosted and demo latency is adjusted by x2 plus 0.15s to approximate production load, so those numbers are not comparable with the raw latency their authors publish.
| Attribute | Jev 1.13.0 | Imajev-4B |
|---|---|---|
| What it is | Hosted System One model from TypeSafe, version jev-1.13.0 | Open LoRA adapter and decision readout on Qwen3.5-4B; also 2B and 9B |
| Licence | Proprietary, hosted | Apache-2.0 adapters and code; Apache-2.0 Qwen base |
| Request format | POST /v1/systemone | The same contract, plus images, unknown_probability and abstained |
| Inputs | Text only: string, JSON object or array | Text, plus up to two images per request |
| Question types | Noul, choice, score — many per call, answered in parallel | Noul, choice, score — 1 to 8 questions and 2 to 254 options per request |
| Context | 64k tokens per request | State up to 32 KB, about 8k tokens |
| Abstaining | Not documented as a separate output | A trained unknown option on every answer |
| Calibration | Post-trained with RLCD; JevBench 76.3 | One fitted temperature per size, 1.305 for the 4B; JevBench 80.4 |
| Languages | Best in English; other languages work with lower accuracy | English only |
| Hardware | None — it is an API call | Apple silicon through MLX or one CUDA GPU; 9.3 GB base weights for the 4B |
Where Jev and Imajev-4B really differ, and why the bars above land where they do.
Imajev-4B is the new first place on JevBench. Version 1.4.2.2, published on 27 September, scores it 67.4 against Jev’s 63.3: #1 of 91, with Jev #4. It is also first when accuracy is weighted at 60%, where Jev is third, and it has the best calibration score in the top ten, 80.4 against Jev’s 76.3. On speed it scores 90.6 against 83.3, from a 40ms raw median on the benchmark’s GPU.
It is not ahead everywhere. Jev keeps the highest intelligence score in the top ten, 53.1 against 52.2, and JevBench’s intelligence-only view puts Jev first and Imajev-4B third. Jev answers 94.5% of the judge tier against 89.0%. On the sealed set the two are level, 36.7% against 37.0%. For the hard tier the board publishes Imajev-4B’s two halves but no combined figure: 72.1% of the 111 public hard decisions and 75.2% of the 109 held-out ones, which works out to 73.6% of all 220 against Jev’s 74.1%.
The larger difference is what it can look at. Jev reads text only; Imajev takes up to two photos in the same request, such as a reference and a target, and checks them against the record you send. Every answer also carries a trained probability that the evidence cannot settle the question, so your code can hand those cases to a person. JevBench is text-only and measures neither.
Imajev-4B, the open image-capable decision model, by Mohit Garg.
Imajev is a family of open decision models by Mohit Garg in three sizes, 2B, 4B and 9B, each a LoRA adapter and a 255-code decision readout on Qwen3.5. The 4B, the recommended default, keeps Qwen3.5-4B’s vision tower frozen. It reads the state, up to two images and each question, then takes the probability of each option from one position without generating text. The adapters and code are Apache-2.0, as is the Qwen base.
Its server accepts the same /v1/systemone request and response as TypeSafe’s Jev and adds images, an unknown probability and an abstained flag; the README says text-only Jev requests work unchanged. Per request it takes up to two images, a state up to 32 KB, one to eight questions and 2 to 254 options. It runs on Apple silicon through MLX or on a CUDA GPU, and the 4B needs a 9.3 GB base-model download. The author measures about 0.1s per question raw on one H100, and about 0.35s as shipped, averaging four option orders.
It was trained on about a million decisions in four stages, labelled by people or by open-weight teacher models. The author states that no Jev outputs, no paid-API outputs and no JevBench items were used, and that JevBench items were screened out with an 8-gram check. JevBench measured the released 4B adapter with one option order and the shipped calibration file, rather than the four-order averaging the README uses for its own figures.
Imajev-4B is ahead on six of the ten sealed families: trap and adversarial items, 58.3% against 41.7%; hard judging, 43.9% against 34.1%; ambiguous questions where abstaining is right, 40.5% against 29.7%; safety judgments, 43.8% against 37.5%; long policies, 32.5% against 27.5%; and dates and numbers, 30.4% against 28.6%. The two tie on trade-offs at 38.5%.
Jev leads by wide margins on the other three: probability questions, 50.0% against 25.0%; paraphrase robustness, 64.3% against 50.0%; and multi-hop lookups, 44.7% against 34.2%. Every system of this kind scores low on the sealed set overall, not far above the 29.3% chance level, so treat these as a guide to what to test rather than a verdict.
The README is candid. On the public hard split alone, two other open models, JevK5 and Eikos-4B, are ahead of every Imajev size. Without its calibration file the model is over-confident on hard items. An empty field can be read as “no” instead of unknown, and English is the only supported language.
Its own image benchmark, ImajevBench, uses AI-generated images that have not yet had a human audit, and it was one input to choosing the released checkpoints. Benchmark Heaven’s separate Image JevBench is an independent check on the photos; the Imajev-4B vs Jev-Omni page covers it.
Imajev-4B is close to Jev on text and does something Jev cannot with photos, so the useful test is your own cases. Run the text ones on Jev here with no setup and keep the answers. If your decisions need a photo, this site also hosts Jev-Omni, the other open decision model that reads images.
On JevBench’s composite, yes: #1 at 67.4 against Jev’s #4 at 63.3, with better calibration and speed. Jev keeps a slightly higher intelligence score, 53.1 against 52.2, and a clear lead on the judge tier, 94.5% against 89.0%. On the sealed set they are level.
Yes. It takes up to two images per request, for example a reference and a target, alongside the state and questions. JevBench, which ranks it first, is a text-only benchmark and does not test that.
No. It is an independent Apache-2.0 project by Mohit Garg that mirrors Jev’s request contract. Its author states that no Jev outputs were used in training.
Its README says text-only requests written for Jev’s /v1/systemone work unchanged, and the response adds unknown_probability and abstained fields. Check your state size against its 32 KB limit first.
Both are open decision models that read images. On the text JevBench, Imajev-4B is #1 and Jev-Omni #11; on Benchmark Heaven’s separate Image JevBench, Jev-Omni is #1 and Imajev-4B #11. The Imajev-4B vs Jev-Omni page puts them side by side.
Every number on these pages is quoted from a published source and was read on September 28, 2026.
Imajev-4B and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on September 28, 2026 and can change.