Jev AI

Two open image decision models · updated September 28, 2026

Imajev-4B vs Jev-Omni

Both read a photo, take your typed questions and return a probability for every option instead of prose. Each is first on a different leaderboard: Imajev-4B on the text JevBench, Jev-Omni on Image JevBench. Here is why, and which one fits your decisions.

Jev-Omni is hosted in our Playground in beta. Imajev-4B is not offered on this site, and neither model is available through our public API.

QWEN3.5 FAMILY

Imajev-4B

Photos checked against your record, with a trained “can’t tell”.

4B
Two photos Can’t tell Typed questions
Imajev-4B against Jev mohit67890/imajev · also 2B and 9B
GEMMA 4 BACKBONE

Jev-Omni

One classifier for text, images, audio and video.

12B
Images Audio Video
Explore Jev-Omni akhilaaa3/Jev-Omni · hosted here

Measured by the same third party

Two leaderboards, opposite orders

Benchmark Heaven runs both boards. On text, Imajev-4B is ahead on every measure shown. On images it answers more items correctly, but Jev-Omni is faster and its composite, which also weighs a cost estimate, comes out first.

MeasureImajev-4BJev-Omni
JevBench v1.4.2.2 · text, 842 decisions
Rank and composite#1 of 91 · 67.4#11 of 91 · 51.3
Intelligence52.246.8
Calibration80.464.1
Sealed accuracy37.0%32.1%
Image JevBench v0.1.2 · images, 684 decisions
Rank and composite#11 of 49 · 65.72#1 of 49 · 73.10
Intelligence73.863.9
Calibration89.289.9
Speed86.289.7
Median response0.132s0.076s

Imajev-4B on Image JevBench

Public image items81.6%
186 of 228 correct
Sealed image items83.1%
379 of 456 correct

Jev-Omni on Image JevBench

Public image items67.1%
153 of 228 correct
Sealed image items80.7%
368 of 456 correct

Image JevBench notes that 333 of its sealed items are its own synthetic images and renders, which most systems find easier than the older real-source items. Imajev-2B, the smaller size, ranks #5 on the same board at 68.72, above the 4B, because it is cheaper and faster to run.

Read the test, not just the number

The Imajev author’s own image test

ImajevBench is the Imajev author’s benchmark: 279 test questions about photos, records and text, 21 of which have no answer the evidence supports. It is a preview with AI-generated images and no human audit yet.

ACCURACY
83.9% vs 78.5%

Imajev-4B against Jev-Omni, run by the Imajev author with Jev-Omni’s own predict(). The author’s significance test, run on an earlier Imajev release, found a gap of this size not significant.

CAN’T TELL
18 vs 0 of 21

The gap comes from the “can’t tell” questions, which Jev-Omni has no output for. On answerable questions the author found Jev-Omni slightly ahead of an earlier Imajev release, and better calibrated.

What actually differs

Both return decision probabilities instead of generated explanations. Their inputs, limits and serving paths are not the same.

CapabilityImajev-4BJev-Omni
BackboneQwen3.5-4B with a LoRA adapter and decision readout; also 2B and 9BGemma 4 12B IT, fine-tuned and merged
ImagesUp to two per request, such as a reference and a targetOne per request
Audio and videoNot supportedAudio up to 30 seconds; video as 16 sampled frames
Can’t tellA trained unknown probability and an abstained flag on every answerNo abstain output
Request formatTypeSafe-style /v1/systemone, plus imagesIts own Python loader and predict(); hosted here in the Playground
Options2 to 254 per question, 1 to 8 questions per requestHead accepts 256; best supported at 20 or fewer
HardwareApple silicon through MLX or one CUDA GPU; 9.3 GB base weightsA CUDA GPU; about 24 GB of BF16 weights
LicenceApache-2.0 adapters and code on an Apache-2.0 baseApache-2.0 weights following Gemma 4; dataset rights separate
Try itThe author’s Hugging Face demo, or run it yourselfHosted in this site’s Playground (sign in), or run it yourself

Which should you evaluate?

Start with Imajev-4B when…

  • You compare a photo with a record: a listing, a return, a part against a known-good one.
  • You want the model to say it cannot tell, and send those cases to a person.
  • Many of your decisions are text-only, where it leads the text board.
  • You need to run it locally, including on a Mac.
Imajev-4B against Jev, in detail →

Start with Jev-Omni when…

  • Your decisions involve audio or video, not only photos.
  • Response time on image decisions matters most.
  • You want to try it right now without a GPU, in this site’s Playground.
Jev-Omni capabilities and setup →

Imajev-4B vs Jev-Omni FAQ

Which is better on images, Imajev-4B or Jev-Omni?

It depends on what you count. On Benchmark Heaven’s Image JevBench, Imajev-4B answers more items correctly, 81.6% of the public ones against 67.1% and 83.1% of the sealed ones against 80.7%, but Jev-Omni ranks first overall because it is faster and cheaper per decision. On the Imajev author’s own ImajevBench, Imajev-4B scores 83.9% against 78.5%.

Which is better on text?

Imajev-4B. On the text JevBench v1.4.2.2 it is #1 of 91 at 67.4 and Jev-Omni is #11 at 51.3, and Imajev-4B is ahead on intelligence, calibration and the sealed set.

Can either model say it cannot tell?

Imajev-4B can: every answer carries a trained unknown probability and an abstained flag. Jev-Omni has no abstain output, so on questions whose honest answer is “can’t tell” it always picks one of the options.

Can I try them without a GPU?

Jev-Omni is hosted in this site’s Playground: sign in and run text, image, audio or video decisions. Imajev-4B has a demo on Hugging Face Spaces run by its author, and it also runs locally on Apple silicon.

Are these made by TypeSafe?

No. Imajev is an independent project by Mohit Garg and Jev-Omni an independent model by akhilaaa3. Both borrow Jev’s typed-decision idea; TypeSafe’s Jev itself reads text only.

Sources read September 28, 2026: JevBench v1.4.2.2, Image JevBench v0.1.2, Imajev repository and Jev-Omni model card. Imajev and Jev-Omni belong to their respective authors and are not affiliated with Jev AI or TypeSafe.