TypeSafe
Jev
Keep an established Jev workflow when its measured accuracy and thresholds meet your requirements. Cloudflare's comparison still gives Jev the lead on some tasks.
Model comparison · Updated October 2, 2026
Jev, Clef and Clef Flash turn state and typed questions into decision probabilities. Compare their capabilities and Cloudflare's published evaluation results before choosing a model for your workflow.
TypeSafe
Keep an established Jev workflow when its measured accuracy and thresholds meet your requirements. Cloudflare's comparison still gives Jev the lead on some tasks.
Cloudflare · 27B
Evaluate the larger Cloudflare model for classification quality, a larger context window or an open-weight deployment.
Cloudflare · 9B
Evaluate the smaller model for latency-sensitive decisions. Its reported speed comes with task-dependent quality differences.
| Capability | Jev | Clef | Clef Flash |
|---|---|---|---|
| Developer | TypeSafe | Cloudflare | Cloudflare |
| Typed outputs | Yes/no, choice, score | Yes/no, choice, score | Yes/no, choice, score |
| Documented model context | 32K (Cloudflare launch comparison) | 65,536 tokens | 65,536 tokens |
| Model-level vision | Text-only native Jev | Embedded images | Embedded images |
| Inputs in this Jev AI integration | Text and JSON | Text and JSON; images not enabled | Text and JSON; images not enabled |
| Hosting | Hosted Jev API | Workers AI or self-hosted weights | Workers AI or self-hosted weights |
| Jev AI API selector | jev-latest / jev-preview / jev-1.13.0 | clef | clef-flash |
The documented context window is a model limit, not a promise that this site's input limit is the same. Jev AI's playground and API limits apply. Cloudflare may truncate long text to fit its model context.
The following is the complete ten-task table from Cloudflare's October 1 launch article, restricted to the three models compared here. Higher is better within each row. Metrics differ between rows; these values are not averaged into a new overall score.
Vendor-reported results, not measurements taken through Jev AI. This is a different evaluation from the JevBench v1.5.0 snapshot on our other comparison pages. Scores and rankings should not be mixed.
| Task / metric | Jev | Clef | Clef Flash |
|---|---|---|---|
| BFCLCase exact | 95.75 | 98.47 | 98.76 |
| ToolRetnDCG@10 | 65.28 | 69.19 | 66.43 |
| API-BankAccuracy | 88.19 | 91.93 | 93.11 |
| Home appliancesCase exact | 52.27 | 82.95 | 97.73 |
| When2CallAccuracy | 80.97 | 72.37 | 65.58 |
| BANKING77Macro-F1 | 79.74 | 94.20 | 90.93 |
| CLINC150+OOSMacro-F1 | 89.27 | 97.43 | 66.77 |
| BRIGHTnDCG@10 | 47.52 | 45.91 | 39.26 |
| Amazon ESCIMacro-F1 | 55.21 | 57.48 | 57.39 |
| PhishNChipsAccuracy | 62.55 | 79.60 | 75.05 |
Clef leads Jev on eight of these ten rows. Jev leads on When2Call and BRIGHT. Clef Flash also trails Jev on CLINC150+OOS, which is a reason to test out-of-scope detection before replacing a classifier.
Cloudflare also reports these four results against TypeSafe's workflow evaluation suite. The best choice changes with the workflow.
| Workflow | Jev | Clef | Clef Flash |
|---|---|---|---|
| Invoice processing | 61.8 | 64.7 | 57.1 |
| Customer service | 76.0 | 76.3 | 77.0 |
| Security incidents | 61.7 | 62.9 | 61.7 |
| Agent trace observability | 71.6 | 68.5 | 69.8 |
| Measurement | Jev | Clef | Clef Flash |
|---|---|---|---|
| Median latency | 524.1 | 209.3 | 38.8 |
| p95 latency | 536.0 | 238.6 | 122.4 |
These figures are not an end-to-end Jev AI latency guarantee. Network distance, input length, question count, queues and the serving path affect your result.
Cloudflare lists Clef at $0.24 per million input tokens and Clef Flash at $0.09 per million input tokens. Those are upstream Workers AI prices, not Jev AI plan prices. Requests through this site use the existing Jev AI credit and token billing. View Jev AI pricing.
No. Clef is a separate Cloudflare model family. It uses a compatible state-and-questions interface, so the request structure is familiar, but model behavior, probabilities and limits can differ.
No. Cloudflare reports stronger Clef results on several classification and tool benchmarks. Its own table also shows Jev ahead on When2Call and BRIGHT, and on agent trace observability in the workflow evaluation. Evaluate the tasks and error types that matter to your application.
Yes, for models configured in this environment. Send the selected model to POST /api/v1/systemone. Check authenticated GET /api/v1/models first. Clef requests are sent to Cloudflare and are never silently replaced with Jev.
This integration currently accepts text and JSON only. Cloudflare documents embedded-image input on Workers AI, but that input path is not enabled through Jev AI. Native Jev and the separate Jev-Omni model should not be confused.
Reuse your question schema as a starting point, then evaluate the new model on held-out examples. A threshold chosen for one model can produce a different false-positive or false-negative rate on another.
Reviewed October 2, 2026. Jev AI is an independent integration, not the official site of TypeSafe or Cloudflare.