Jev AI

Jev input guide · updated October 4, 2026

Is Jev multimodal?

No. Native Jev reads text only: a string, a JSON object or an array. It cannot see an image, hear audio or watch video. On Jev AI, two other models in the same playground can: Clef answers typed questions about images, and Jev-Omni about images, audio and video.

What each model on Jev AI can read

All five use the same editor and return the same kind of answer: probabilities, not prose.

Jev facts are from TypeSafe’s public documentation as of September 18, 2026. Clef, Jev-Omni and Laya are separate models from other developers; Jev AI never substitutes one for another. Availability can vary, so check the model menu in the playground.

Images: use Clef

The closest thing to “Jev with vision” for typed decisions about pictures.

Clef and Clef Flash are Cloudflare’s decision models. They accept the same state-and-questions request as Jev and add a vision encoder, so a question can be about a screenshot, a product photo or a scanned document.

  • Up to 4 PNG, JPEG or WebP images per request.
  • Each image up to 4 MiB and 16 megapixels; 8 MiB in total.
  • Images must be embedded in the request. Remote image URLs are not accepted.
  • Works in the playground and through the API, with all three question types.

Clef models, limits and API example · Jev vs Clef

GPT-6 Luna, OpenAI’s Decisions API model, accepts the same embedded images with the same limits on this site.

Audio and video: use Jev-Omni

An independent open-weights classifier, hosted in the Jev AI playground.

Jev-Omni is built on Gemma 4 12B by akhilaaa3. It is not a TypeSafe model. It answers one Choice question about a text, an image, an audio clip or a video, and returns a probability for each option.

  • One JPG, PNG or WebP image, an audio clip of up to 30 seconds, or a video of up to 60 seconds sampled as 16 frames.
  • One file per run, up to 6 MiB.
  • One Choice question with 2 to 20 options. Yes/no and score questions are not available.
  • Playground only. Public API access is not included.

Jev-Omni benchmarks, hardware and limits · OmniJev vs Jev-Omni

PDFs, screenshots and scanned documents

Decide whether the content is text or a picture of text.

  1. 01

    A PDF with a text layer

    PDF upload is not supported. Extract the text, keep the passages that matter and send them as the state to Jev. Filtering first helps: accuracy falls when the state is full of unrelated text.

  2. 02

    A scan or photo of a page

    Send the page as an image to Clef, or run your own OCR and give the resulting text to Jev.

  3. 03

    A screenshot

    Attach it to a Clef request and ask your typed questions about it directly.

  4. 04

    Mixed records

    Put structured fields in the state as JSON. Jev reads a JSON object or array as well as plain text.

Self-hosted models that read images

Open projects that answer Jev-style questions about pictures on your own hardware. Jev AI does not host these.

Each linked page cites the project’s own repository or model card; JEMM links to its repository because it has no independent measurement to cite yet. For hardware and licences, see Jev alternatives you can run locally.

Is Jev multimodal? FAQ

Is Jev multimodal?

No. TypeSafe’s Jev accepts text only: a string, a JSON object or an array. It does not read images, audio or video. On Jev AI you can switch the playground to Clef for images or to Jev-Omni for images, audio and video.

Can Jev read images?

Native Jev cannot. To ask a typed question about an image on Jev AI, select Clef or Clef Flash, which accept up to 4 PNG, JPEG or WebP images per request, or Jev-Omni, which reads one image per run.

Does Jev have vision?

Jev 1.13 has no vision input. Clef and Clef Flash include a vision encoder, and Jev-Omni is a multimodal classifier. Both are separate models from different developers that answer Jev-style typed questions.

Is Jev-Omni made by TypeSafe?

No. Jev-Omni is an independent open-weights model by akhilaaa3, built on Gemma 4 12B. It borrows the Jev name and the typed-decision idea but shares no weights with TypeSafe’s Jev.

Can I send images through the Jev API?

Yes, when the model is Clef or Clef Flash: add an images array of embedded images to the request. Remote image URLs are not accepted. Requests to jev-latest must be text. Jev-Omni is available in the playground only, not through the public API.

Can Jev read a PDF?

PDF upload is not supported. Convert the PDF to text and send the relevant passages as the state. For a scanned page with no text layer, send an image of the page to Clef instead.

Can Jev classify audio or video?

Native Jev cannot. In the Jev AI playground, Jev-Omni accepts an audio clip of up to 30 seconds or a video of up to 60 seconds, sampled as 16 frames, in a file of up to 6 MiB, and answers one Choice question about it.

Text decisions? Stay with Jev

Native Jev has the lowest credit cost of the models here. Switch only when the input is not text.

Open the playground

Clef, Jev-Omni, Laya and every other model named here belong to their respective owners. Jev AI is independently operated and is not TypeSafe’s official site. Limits describe Jev AI as of October 4, 2026.