Jev input guide · updated October 4, 2026
Is Jev multimodal?
No. Native Jev reads text only: a string, a JSON object or an array. It cannot see an image, hear audio or watch video. On Jev AI, two other models in the same playground can: Clef answers typed questions about images, and Jev-Omni about images, audio and video.
What each model on Jev AI can read
All five use the same editor and return the same kind of answer: probabilities, not prose.
| Model | Text | Images | Audio | Video | Questions | Public API | Credits per run |
|---|---|---|---|---|---|---|---|
| JevTypeSafe | Yes | No | No | No | Up to 8 per playground run, 64 per API request | Yes | 1 |
| ClefCloudflare | Yes | Up to 4 per request | No | No | Up to 8 per playground run, 64 per API request | Yes, with an images array | 6 |
| Clef FlashCloudflare | Yes | Up to 4 per request | No | No | Up to 8 per playground run, 64 per API request | Yes, with an images array | 3 |
| GPT-6 LunaOpenAI | Yes | Up to 4 per request | No | No | Up to 8 per playground run, 64 per API request | Yes, with an images array | 3 |
| Jev-Omniakhilaaa3 | Yes | One per run | Up to 30 seconds | Up to 60 seconds, 16 frames sampled | One Choice question per run | No, playground only | 4 |
| LayaConvai Innovations | Yes, short text | No | No | No | All three types; 512 or 1,024 tokens per question | Yes | 2 |
Jev facts are from TypeSafe’s public documentation as of September 18, 2026. Clef, Jev-Omni and Laya are separate models from other developers; Jev AI never substitutes one for another. Availability can vary, so check the model menu in the playground.
Images: use Clef
The closest thing to “Jev with vision” for typed decisions about pictures.
Clef and Clef Flash are Cloudflare’s decision models. They accept the same state-and-questions request as Jev and add a vision encoder, so a question can be about a screenshot, a product photo or a scanned document.
- Up to 4 PNG, JPEG or WebP images per request.
- Each image up to 4 MiB and 16 megapixels; 8 MiB in total.
- Images must be embedded in the request. Remote image URLs are not accepted.
- Works in the playground and through the API, with all three question types.
Clef models, limits and API example · Jev vs Clef
GPT-6 Luna, OpenAI’s Decisions API model, accepts the same embedded images with the same limits on this site.
Audio and video: use Jev-Omni
An independent open-weights classifier, hosted in the Jev AI playground.
Jev-Omni is built on Gemma 4 12B by akhilaaa3. It is not a TypeSafe model. It answers one Choice question about a text, an image, an audio clip or a video, and returns a probability for each option.
- One JPG, PNG or WebP image, an audio clip of up to 30 seconds, or a video of up to 60 seconds sampled as 16 frames.
- One file per run, up to 6 MiB.
- One Choice question with 2 to 20 options. Yes/no and score questions are not available.
- Playground only. Public API access is not included.
Jev-Omni benchmarks, hardware and limits · OmniJev vs Jev-Omni
PDFs, screenshots and scanned documents
Decide whether the content is text or a picture of text.
- 01
A PDF with a text layer
PDF upload is not supported. Extract the text, keep the passages that matter and send them as the state to Jev. Filtering first helps: accuracy falls when the state is full of unrelated text.
- 02
A scan or photo of a page
Send the page as an image to Clef, or run your own OCR and give the resulting text to Jev.
- 03
A screenshot
Attach it to a Clef request and ask your typed questions about it directly.
- 04
Mixed records
Put structured fields in the state as JSON. Jev reads a JSON object or array as well as plain text.
Self-hosted models that read images
Open projects that answer Jev-style questions about pictures on your own hardware. Jev AI does not host these.
| Model | What it reads |
|---|---|
| Imajev-4B | Text, plus up to two photos per request |
| Winnow-12B | Text, plus images with its optional vision projector |
| OpenJev (razorback16) | Text, plus up to 8 images per request |
| djev | Text, images, image options and live camera frames |
| OmniJev | Images and screenshots; video as 16 timestamped frames |
| JEMM ↗ | Text, plus up to 4 screenshots per request. A Qwen3.8-27B decision model under Apache-2.0 that needs a 64 GB GPU; not measured on JevBench |
Each linked page cites the project’s own repository or model card; JEMM links to its repository because it has no independent measurement to cite yet. For hardware and licences, see Jev alternatives you can run locally.
Is Jev multimodal? FAQ
Is Jev multimodal?
No. TypeSafe’s Jev accepts text only: a string, a JSON object or an array. It does not read images, audio or video. On Jev AI you can switch the playground to Clef for images or to Jev-Omni for images, audio and video.
Can Jev read images?
Native Jev cannot. To ask a typed question about an image on Jev AI, select Clef or Clef Flash, which accept up to 4 PNG, JPEG or WebP images per request, or Jev-Omni, which reads one image per run.
Does Jev have vision?
Jev 1.13 has no vision input. Clef and Clef Flash include a vision encoder, and Jev-Omni is a multimodal classifier. Both are separate models from different developers that answer Jev-style typed questions.
Is Jev-Omni made by TypeSafe?
No. Jev-Omni is an independent open-weights model by akhilaaa3, built on Gemma 4 12B. It borrows the Jev name and the typed-decision idea but shares no weights with TypeSafe’s Jev.
Can I send images through the Jev API?
Yes, when the model is Clef or Clef Flash: add an images array of embedded images to the request. Remote image URLs are not accepted. Requests to jev-latest must be text. Jev-Omni is available in the playground only, not through the public API.
Can Jev read a PDF?
PDF upload is not supported. Convert the PDF to text and send the relevant passages as the state. For a scanned page with no text layer, send an image of the page to Clef instead.
Can Jev classify audio or video?
Native Jev cannot. In the Jev AI playground, Jev-Omni accepts an audio clip of up to 30 seconds or a video of up to 60 seconds, sampled as 16 frames, in a file of up to 6 MiB, and answers one Choice question about it.
Text decisions? Stay with Jev
Native Jev has the lowest credit cost of the models here. Switch only when the input is not text.
Clef, Jev-Omni, Laya and every other model named here belong to their respective owners. Jev AI is independently operated and is not TypeSafe’s official site. Limits describe Jev AI as of October 4, 2026.