Jev AI

Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs Imajev-4B

The new first place on JevBench, from a 4B model that also reads photos.

On this site

Jev 1.13.0

TypeSafe · System One model

Fourth overall, and still the highest intelligence score in the top ten.

63.3JevBench v1.4.2.2 rank #4

vs

Alternative

Imajev-4B

Mohit Garg · Open Qwen3.5-4B decision model that also reads photos

First of 91, with the best calibration in the top ten.

67.4JevBench v1.4.2.2 rank #1

  1. JevHosted API, generally availableWhere it runsImajev-4BSelf-hosted on a Mac or one GPU
  2. JevText onlyInputsImajev-4BText, plus up to two photos per request
  3. JevNot documented as a separate outputCan’t tellImajev-4BA trained unknown probability on every answer
  4. Jev94.5% correctEvaluation-style casesImajev-4B89.0% correct
  5. Jev36.7% correctSealed decisionsImajev-4B37.0% correct
  6. Jev76.3 on JevBenchCalibrationImajev-4B80.4 on JevBench
  7. Jev64k tokens per requestInput sizeImajev-4BState up to 32 KB, about 8k tokens

JevBench v1.4.2.2, measured the same way

One benchmark measured all 95 systems under one method and ranked 91, so these bars are comparable with each other in a way that vendor-published figures are not. The full board and how to read it are on the JevBench results page.

Composite score

#1 Imajev-4B67.4
#2 Plumb-4B65.8
#3 decider-4b64.1
#4 Jev63.3
#5 JevK562.0
#6 Cygnet61.8
#7 Hopper59.4
#8 Winnow55.6
#10 djev52.2
#13 SemIf47.7
#29 OpenJev36.9
#43 Laya30.3

The top four finish within 4.1 points of each other, and then the board falls away sharply. The composite hides where systems actually differ, so the two charts below break it apart.

Capability, higher is better

IntelligenceHow often it picks the right answer

Jev53.1
Imajev-4B52.2

CalibrationWhether 0.8 really means about 80%

Jev76.3
Imajev-4B80.4

SpeedMeasured response time

Jev83.3
Imajev-4B90.6

Accuracy by how hard the decision is

Easy72 straightforward cases

Jev100%
Imajev-4B100%

Standard96 everyday cases

Jev99%
Imajev-4B99%

Judge146 evaluation-style calls

Jev94.5%
Imajev-4B89%

Hard220 genuinely ambiguous cases

Jev74.1%
Imajev-4B—

Sealed308 private cases, new in v1.4; chance is 29.3%

Jev36.7%
Imajev-4B37%

Easy and standard decisions separate almost nothing. The hard tier, highlighted, is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure for that tier.

One thing these charts leave out: the composite score above also weighs a fourth axis, cost. We do not reproduce it here; the full table is at the source below.

Every system on the board

The capability scores are charted above; this is the ranking and the deployment detail behind them. Jev and Imajev-4B are highlighted.

ModelRankScoreMeasured latencyRuns on
Imajev-4B (Mohit Garg, Qwen3.5-4B LoRA)#167.40.04s raw, 0.23s adjustedSelf-hosted GPU
Plumb-4B (crh225, JevK5 v0.2 + LoRA)#265.8See the source tableSelf-hosted GPU
decider-4b v2 (Mapika, Qwen3.5-4B)#364.10.02s raw, 0.18s adjustedSelf-hosted, RTX PRO 6000
Jev 1.13.0 (TypeSafe)#463.30.65s median / 0.72s p95Hosted production API
JevK5 v0.2.0 (allebee, Qwen3.5-4B)#562.0Not re-measured in v1.4Self-hosted, RunPod GPU
Cygnet (blockbrain, frozen Gemma 4 12B)#661.80.04s raw, 0.22s adjustedSelf-hosted, RTX PRO 6000
Hopper (HopitAI, Qwen3.5-4B LoRA)#759.40.13s raw, 0.41s adjustedSelf-hosted, RTX A6000
Winnow-12B Q8 (EldanRing, Gemma 4 12B)#855.60.23s raw, 0.60s adjustedSelf-hosted, RTX 4090
djev (Maisa, DiffusionGemma)#1052.20.24s median / 0.31s p95Hosted API, preview
SemIf (Qwen3.5-4B)#1347.70.20s raw, 0.55s adjustedSelf-hosted, RunPod GPU
OpenJev (razorback16)#2936.90.24s raw, 0.63s adjustedSelf-hosted, RunPod GPU
Laya (ModernBERT-large)#4330.30.79s raw, 1.72s adjustedSelf-hosted, CPU

JevBench v1.4.2.2, published on 27 September 2026, lists 95 systems and ranks 91 of them. Each is scored on the same 534 public decisions (72 easy, 96 standard, 146 judge, 220 hard) plus 308 sealed decisions whose text stays private, one request at a time. The composite is a harmonic mean of four equally weighted axes (intelligence, calibration, speed and cost), so one weak axis pulls it down hard. Sealed decisions make up 20% of intelligence, and a system that does more than 25 points better on public items than on sealed ones loses intelligence in proportion. Every Jev-style system scores far lower on the sealed set than on the public one, Jev included (36.7% against a 29.3% chance level). The scoring formulas have not changed since v1.4.0 and no earlier measurement has been redone; 19 systems were added over four releases, which is why ranks moved while scores did not. Self-hosted and demo latency is adjusted by x2 plus 0.15s to approximate production load, so those numbers are not comparable with the raw latency their authors publish.

Jev and Imajev-4B feature by feature

AttributeJev 1.13.0Imajev-4B
What it isHosted System One model from TypeSafe, version jev-1.13.0Open LoRA adapter and decision readout on Qwen3.5-4B; also 2B and 9B
LicenceProprietary, hostedApache-2.0 adapters and code; Apache-2.0 Qwen base
Request formatPOST /v1/systemoneThe same contract, plus images, unknown_probability and abstained
InputsText only: string, JSON object or arrayText, plus up to two images per request
Question typesNoul, choice, score — many per call, answered in parallelNoul, choice, score — 1 to 8 questions and 2 to 254 options per request
Context64k tokens per requestState up to 32 KB, about 8k tokens
AbstainingNot documented as a separate outputA trained unknown option on every answer
CalibrationPost-trained with RLCD; JevBench 76.3One fitted temperature per size, 1.305 for the 4B; JevBench 80.4
LanguagesBest in English; other languages work with lower accuracyEnglish only
HardwareNone — it is an API callApple silicon through MLX or one CUDA GPU; 9.3 GB base weights for the 4B

What each one is better at

Where Jev wins

  • Intelligence 53.1 against 52.2, the highest in the top ten, and first on JevBench’s intelligence-only view, where Imajev-4B is third.
  • Judge tier 94.5% against 89.0%.
  • Far ahead on the sealed set’s probability questions, 50.0% against 25.0%, and ahead on paraphrase and multi-hop families.
  • 64k tokens per request against a state of about 8k tokens.
  • Nothing to download or serve, and a hosted version with published rate limits.

Where Imajev-4B wins

  • #1 of 91 in JevBench v1.4.2.2 at 67.4 against Jev’s 63.3, and first again with accuracy or cost weighted more heavily.
  • Calibration 80.4 against 76.3, the best in the top ten, and speed 90.6 against 83.3.
  • Reads up to two photos per request and checks them against your record. Jev reads text only.
  • A trained unknown on every answer, so cases the evidence cannot settle can go to a person.
  • Ahead of Jev on six of the ten sealed families, including trap and adversarial items, 58.3% against 41.7%.

The analysis

Where Jev and Imajev-4B really differ, and why the bars above land where they do.

Imajev-4B is the new first place on JevBench. Version 1.4.2.2, published on 27 September, scores it 67.4 against Jev’s 63.3: #1 of 91, with Jev #4. It is also first when accuracy is weighted at 60%, where Jev is third, and it has the best calibration score in the top ten, 80.4 against Jev’s 76.3. On speed it scores 90.6 against 83.3, from a 40ms raw median on the benchmark’s GPU.

It is not ahead everywhere. Jev keeps the highest intelligence score in the top ten, 53.1 against 52.2, and JevBench’s intelligence-only view puts Jev first and Imajev-4B third. Jev answers 94.5% of the judge tier against 89.0%. On the sealed set the two are level, 36.7% against 37.0%. For the hard tier the board publishes Imajev-4B’s two halves but no combined figure: 72.1% of the 111 public hard decisions and 75.2% of the 109 held-out ones, which works out to 73.6% of all 220 against Jev’s 74.1%.

The larger difference is what it can look at. Jev reads text only; Imajev takes up to two photos in the same request, such as a reference and a target, and checks them against the record you send. Every answer also carries a trained probability that the evidence cannot settle the question, so your code can hand those cases to a person. JevBench is text-only and measures neither.

What Imajev-4B actually is

Imajev-4B, the open image-capable decision model, by Mohit Garg.

Imajev is a family of open decision models by Mohit Garg in three sizes, 2B, 4B and 9B, each a LoRA adapter and a 255-code decision readout on Qwen3.5. The 4B, the recommended default, keeps Qwen3.5-4B’s vision tower frozen. It reads the state, up to two images and each question, then takes the probability of each option from one position without generating text. The adapters and code are Apache-2.0, as is the Qwen base.

Its server accepts the same /v1/systemone request and response as TypeSafe’s Jev and adds images, an unknown probability and an abstained flag; the README says text-only Jev requests work unchanged. Per request it takes up to two images, a state up to 32 KB, one to eight questions and 2 to 254 options. It runs on Apple silicon through MLX or on a CUDA GPU, and the 4B needs a 9.3 GB base-model download. The author measures about 0.1s per question raw on one H100, and about 0.35s as shipped, averaging four option orders.

It was trained on about a million decisions in four stages, labelled by people or by open-weight teacher models. The author states that no Jev outputs, no paid-API outputs and no JevBench items were used, and that JevBench items were screened out with an 8-gram check. JevBench measured the released 4B adapter with one option order and the shipped calibration file, rather than the four-order averaging the README uses for its own figures.

Where each one wins on the sealed set

Imajev-4B is ahead on six of the ten sealed families: trap and adversarial items, 58.3% against 41.7%; hard judging, 43.9% against 34.1%; ambiguous questions where abstaining is right, 40.5% against 29.7%; safety judgments, 43.8% against 37.5%; long policies, 32.5% against 27.5%; and dates and numbers, 30.4% against 28.6%. The two tie on trade-offs at 38.5%.

Jev leads by wide margins on the other three: probability questions, 50.0% against 25.0%; paraphrase robustness, 64.3% against 50.0%; and multi-hop lookups, 44.7% against 34.2%. Every system of this kind scores low on the sealed set overall, not far above the 29.3% chance level, so treat these as a guide to what to test rather than a verdict.

What its author says about its limits

The README is candid. On the public hard split alone, two other open models, JevK5 and Eikos-4B, are ahead of every Imajev size. Without its calibration file the model is over-confident on hard items. An empty field can be read as “no” instead of unknown, and English is the only supported language.

Its own image benchmark, ImajevBench, uses AI-generated images that have not yet had a human audit, and it was one input to choosing the released checkpoints. Benchmark Heaven’s separate Image JevBench is an independent check on the photos; the Imajev-4B vs Jev-Omni page covers it.

So which should you use?

Pick Jev if

  • Your inputs are long text: documents, traces or tickets past about 8k tokens.
  • Your hard cases turn on probabilities, reworded questions or multi-step lookups.
  • You want a hosted, versioned model with nothing to serve.

Pick Imajev-4B if

  • Your decisions depend on a photo: a listing against its picture, a return against what was shipped, a part against a known-good one.
  • You want the model to say it cannot tell, and send those cases to a person.
  • You need open weights that run on a Mac or one GPU inside your own network.

Try Jev on the cases you are actually arguing about

Imajev-4B is close to Jev on text and does something Jev cannot with photos, so the useful test is your own cases. Run the text ones on Jev here with no setup and keep the answers. If your decisions need a photo, this site also hosts Jev-Omni, the other open decision model that reads images.

Jev vs Imajev-4B FAQ

Is Imajev-4B better than Jev?

On JevBench’s composite, yes: #1 at 67.4 against Jev’s #4 at 63.3, with better calibration and speed. Jev keeps a slightly higher intelligence score, 53.1 against 52.2, and a clear lead on the judge tier, 94.5% against 89.0%. On the sealed set they are level.

Can Imajev read images?

Yes. It takes up to two images per request, for example a reference and a target, alongside the state and questions. JevBench, which ranks it first, is a text-only benchmark and does not test that.

Is Imajev made by TypeSafe?

No. It is an independent Apache-2.0 project by Mohit Garg that mirrors Jev’s request contract. Its author states that no Jev outputs were used in training.

Do Jev requests work with Imajev?

Its README says text-only requests written for Jev’s /v1/systemone work unchanged, and the response adds unknown_probability and abstained fields. Check your state size against its 32 KB limit first.

How does Imajev-4B compare with Jev-Omni?

Both are open decision models that read images. On the text JevBench, Imajev-4B is #1 and Jev-Omni #11; on Benchmark Heaven’s separate Image JevBench, Jev-Omni is #1 and Imajev-4B #11. The Imajev-4B vs Jev-Omni page puts them side by side.

Other Jev comparisons

Every number on these pages is quoted from a published source and was read on September 28, 2026.

Sources

Imajev-4B and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on September 28, 2026 and can change.