Jev AI

Model guide · multimodal decision model · published October 11, 2026

Clef-Omni

Cloudflare’s decision model that hears and sees: typed answers about text, photos, audio and video in one call. We tested it on all three against the models Jev AI runs.

Clef-Omni joins Cloudflare’s Clef and Clef-flash as the family’s multimodal member. It takes a state plus up to four images, four audio clips and two videos, and returns a probability for every option of every question, with open Apache-2.0 weights. We measured three things: text decisions against Jev on the same day and route, photo and speech decisions against Jev-Omni, the multimodal model Jev AI self-hosts, and how it behaves through OpenRouter as well as Cloudflare.

100% vs 96%100 photos, 10 options: Clef-Omni against Jev AI’s self-hosted Jev-Omni
91% vs 11%100 short spoken digits, same pair; Jev-Omni answered “four” 99 times
76.0% vs 91.0%text: 300 HaluEval verdicts against Jev, where Jev stays ahead
178 msmedian photo decision on Workers AI from our server; Jev-Omni 768 ms

Clef-Omni at a glance

DeveloperCloudflare
Released9 October 2026, on Workers AI as @cf/cloudflare/clef-omni and on OpenRouter as cloudflare/clef-omni
ArchitectureMixture of experts, 30B parameters with 3B active, post-trained from Qwen3-Omni-30B-A3B-Instruct; a joint schema head scores every option of every question in one forward pass
TrainingFrozen backbone with LoRA adapters; label-smoothed cross-entropy plus a Brier-score calibration term
WeightsApache-2.0 on Hugging Face; Cloudflare tested it on one H200, about 64 GB of GPU memory in bf16
InputsText or JSON state, plus up to 4 images, 4 audio clips (wav or mp3, up to 300 s each) and 2 videos (up to 60 s, sampled at 2 frames per second) as base64 data URLs; remote URLs are refused
Context64,000 tokens on Workers AI, shared by media and text; text state is truncated to fit
Media sizeEach image 64–1,024 tokens (32×32-pixel patches); audio about 780 tokens a minute; video up to about 15,400 tokens a minute
APIJev/System One compatible: state, typed questions, answers with probabilities; images, audio and videos are Clef extensions

Our test: photos and speech

Same items, same question and options for both models, one request at a time from the Jev AI production server. Jev AI production server, one request at a time; Clef-Omni on Cloudflare Workers AI, Jev-Omni on Jev AI’s self-hosted GPU gateway.

TaskItemsClef-Omni accuracyJev-Omni accuracyClef-Omni medianJev-Omni medianCalibration error (Clef · Jev-Omni)
Photos: what is in the picture (10 options)100100%96%178 ms768 ms0.077 · 0.016
Spoken digit, original 8 kHz clips (10 options)10091%11%175 ms789 ms0.161 · 0.171
Spoken digit, resampled to 16 kHz10092%11%139 ms1030 ms0.178 · 0.172
Spoken digit, 16 kHz padded to 2 s20100%45%161 ms4222 ms0.177 · 0.183

What we found

  • Photos: Clef-Omni named the subject of all 100 Imagenette photos. Jev-Omni got 96, with better-calibrated probabilities (error 0.016 against 0.077): Clef-Omni was right every time but under 90% sure on a third of the photos, so its probabilities understate its accuracy here.
  • Speech: Clef-Omni recognised the spoken digit in 91 of 100 clips; its misses were near-sounds such as four heard as zero. Jev-Omni answered “four” for 99 of the 100, which is chance. Resampling to 16 kHz changed nothing (11%); padding the clips to two seconds lifted it to 45% on 20 clips, where Clef-Omni scored 100%.
  • Speed: about 178 ms per photo and 175 ms per clip on Workers AI, four to five times quicker than Jev-Omni through our gateway, which serves it from a single GPU.
  • What this means for Jev AI: for short speech, Clef-Omni is clearly the stronger model today, and it is the obvious hosted candidate for the audio side of our multimodal playground. We have opened an investigation into Jev-Omni’s short-clip audio.

How we tested

  • Photos: the first 10 validation images of each of Imagenette’s 10 ImageNet classes (160 px), one Choice per photo: “What does the photo show?” with the 10 class names.
  • Speech: 10 clips per digit from the Free Spoken Digit Dataset, rotating across its six speakers, 0.16–0.82 seconds each, 8 kHz mono; one Choice: “Which digit is spoken?” with zero to nine.
  • Endpoints: Clef-Omni through Cloudflare’s Workers AI REST API with its images and audio fields; Jev-Omni through the same gateway the Jev AI playground uses.
  • Limits: two small public datasets, closed-set labels, no video. Speech digits are a narrow test of hearing; sounds, music and long recordings may rank the models differently.

Our test: text decisions against Jev

300 HaluEval verdicts and 440 BFCL tool-routing decisions, both models through OpenRouter Decisions API from a laptop in Asia, 4 requests at a time, 2026-10-11.

ModelJudge accuracyAccepts · rejectsJudge ECERouting accuracyRight tool · declinesRouting ECEMedian
Clef-Omni76.0%94.7% · 57.3%0.15786.1%100.0% · 74.6%0.086546 ms
Jev (jev-1.13.0)91.0%94.7% · 87.3%0.05793.6%99.5% · 88.8%0.046623 ms

Clef-Omni accepted correct answers as often as Jev but rejected only 57.3% of hallucinated ones, against 87.3%; it picked the right tool every time one fitted, but declined only 74.6% of requests with no fitting tool, against 88.8%. As with the small open models on our Ollama page, the gap is in saying no. For text-only work, Jev or Cloudflare’s text-focused Clef are the stronger choices.

What Cloudflare reports

From the launch post. Cloudflare’s numbers, not ours; Jev’s column is Cloudflare’s run of Jev.

BenchmarkMetricClef-OmniClefClef-flashJev
BFCLcase exact98.298.4798.7695.75
ToolRetnDCG@1066.669.1966.4365.28
API-Bankaccuracy92.791.9393.1188.19
Home appliancescase exact69.382.9597.7352.27
When2Callaccuracy63.372.3765.5880.97
BANKING77macro-F194.894.290.9379.74
CLINC150+OOSmacro-F197.797.4366.7789.27
BRIGHTnDCG@104245.9139.2647.52
Amazon ESCImacro-F157.857.4857.3955.21
PhishNChipsaccuracy73.279.675.0562.55
TypeSafe workflow evalMetricClef-OmniClefClef-flashJev
Invoice processingExact actions60.264.757.161.8
Invoice processingPrimary action8286.273.383.1
Customer serviceExact actions71.676.37776
Security incidentsExact actions61.762.961.761.7
Agent trace observabilityPrimary action65.868.569.871.6

On Cloudflare’s table Clef-Omni beats Jev on 8 of 10 text benchmarks, led by intent classification (BANKING77, CLINC150), and trails it on When2Call and BRIGHT. On TypeSafe’s workflow evals it sits just below Jev and below the dense Clef. Cloudflare reports about 130 ms median for text-only decisions and about 150 ms for image inputs on Workers AI. Our text results above are closer to the When2Call side: the hard part for Clef-Omni is recognising when the answer is no.

Call it

Workers AI for every modality; OpenRouter for text and images.

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-omni \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "clef-omni",
    "state": "Review the installation: a photo of the unit and a recording of it running.",
    "images": ["data:image/jpeg;base64,<base64-jpeg>"],
    "audio": ["data:audio/wav;base64,<base64-wav>"],
    "questions": {
      "label_visible": { "type": "noul", "instructions": "Is the serial number label visible in the photo?" },
      "sounds_normal": { "type": "noul", "instructions": "Does the unit run smoothly, without rattling or grinding?" }
    }
  }'

Through OpenRouter, images go inside the state array, and the top-level images field is refused:

{
  "model": "cloudflare/clef-omni",
  "state": [
    { "type": "text", "text": "Inspect the attached photo." },
    { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,<base64-jpeg>" } }
  ],
  "questions": { "subject": { "type": "choice", "instructions": "What does the photo show?",
                              "criteria": { "a church": null, "a garbage truck": null, "a parachute": null } } }
}

Audio did not get through OpenRouter in our test: every format we tried was counted as thousands of text tokens and answered “zero”. Media count toward the 64,000-token context; if they overflow it the request fails, otherwise the text state is truncated to fit.

Which model for which input

Clef-Omni

  • Speech or sounds in the state: it heard spoken digits where Jev-Omni could not.
  • Photos, with the fastest decisions we measured.
  • Mixed media in one call: a photo, a recording and a clip answered together.

Jev

  • Text and JSON decisions where rejecting a wrong answer matters: it led Clef-Omni by 30 points there.
  • Agent routing that must decline when no tool fits. Try it in the playground.

Clef-Omni FAQ

What is Clef-Omni?

Cloudflare’s multimodal decision model, released on 9 October 2026. It reads a state made of text, JSON, images, audio or video and answers typed noul, choice and score questions with a probability for every option, in one forward pass and without generating text. It is a 30B mixture-of-experts model with 3B active parameters, post-trained from Qwen3-Omni-30B-A3B-Instruct, with Apache-2.0 weights.

How does Clef-Omni compare with Jev?

They do different jobs. Jev reads text only; Clef-Omni also reads images, audio and video. On text, Jev was stronger in our test the same day through the same route: 91.0% against 76.0% on 300 HaluEval verdicts and 93.6% against 86.1% on 440 tool-routing decisions. Cloudflare’s own table has Clef-Omni ahead on 8 of 10 benchmarks, mostly intent classification and tool calls.

Can Clef-Omni understand speech?

It recognised the spoken digit in 91 of 100 short clips (0.16–0.82 seconds, six speakers) through Cloudflare’s API, and in all 20 clips padded to two seconds. It hears the audio directly; there is no transcription step.

Does Clef-Omni work through OpenRouter?

For text and images, yes. Images go in the state array as image_url parts with a base64 data URL; the top-level images field is refused there. Audio did not pass through OpenRouter in our test on 11 October 2026: every format we tried was counted as text and answered at chance. Use Cloudflare’s Workers AI API for audio and video.

Can I run Clef-Omni myself?

Yes. The weights are Apache-2.0 on Hugging Face with inference code that also exposes the same /v1/systemone request shape. Cloudflare tested it on a single H200; the backbone needs about 64 GB of GPU memory in bf16.

How fast is it?

From our server to Workers AI, a photo decision took 178 ms and an audio decision 175 ms at the median, one request at a time. Cloudflare reports about 130 ms for text and 150 ms for images.

Is Clef-Omni in the Jev AI playground?

Not yet. Jev AI’s playground runs Jev, Clef, Clef Flash, GPT-6 Luna, Mercury Decide, Laya and the self-hosted Jev-Omni. This page measures Clef-Omni directly on Cloudflare.

About this page

Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and self-hosts Jev-Omni, so the media comparison is against our own service. We are not affiliated with Cloudflare or TypeSafe.

How. Facts and vendor tables are from Cloudflare’s announcement, Workers AI documentation and the Hugging Face model card. Our numbers come from the scripts and raw result files in our repository, run on 2026-10-11; the media item list is published with them.

When. Published October 11, 2026. We will rerun audio when Jev-Omni or OpenRouter’s media handling changes.

Sources

Cloudflare, Clef and Clef-Omni are Cloudflare’s; Jev is TypeSafe AI’s. Neither is affiliated with Jev AI.