On this site
Jev 1.13.0
TypeSafe · System One decision model
Typed decisions with calibrated probabilities, in about 0.2 seconds, at $0.042 per million input tokens.
Jev vs a generative model · independent measurements · updated October 7, 2026
A decision model against a generative one: close on labels, far apart on speed, cost and what comes back.
Claude writes, reasons and explains; Jev answers typed questions with a probability and nothing else. People compare them because teams send Claude the same classification, routing and judging prompts Jev was built for. Three independent measurements put the two side by side on exactly those tasks. The short version: Claude Opus is a few points more accurate on hard intent sets, Jev is 5 to 13 times faster and one to two orders of magnitude cheaper, and on narrow decomposed questions Jev matches or beats Haiku. This page quotes the numbers with their sources, then says when each one belongs in your stack.
On this site
TypeSafe · System One decision model
Typed decisions with calibrated probabilities, in about 0.2 seconds, at $0.042 per million input tokens.
vs
Compared with
Anthropic · Generative models: Opus 5.5, Sonnet 5.5, Haiku 4.5
Text, code and reasoning from a 1M-token context, from $1 to $4 per million input tokens and $5 to $20 per million output.
Jev and Claude share no benchmark, so every number here is one party’s own test, quoted with its setup and its caveats.
OpenRouter sorted the 3,080 utterances of the Banking77 test split into 77 banking intents with both models. Jev got one-line criteria per intent and no fine-tuning; Opus ran at temperature zero with reasoning off and a strict JSON response schema. Latency is the client-observed round trip at 8 concurrent requests.
| Jev | Claude | |
|---|---|---|
| Accuracy | 81.0% (95% interval 79.6 to 82.3) | 84.4% (83.1 to 85.6) |
| Macro-F1 | 80.5% | 83.6% |
| Median latency | 175 ms | 2,266 ms |
| p95 latency | 270 ms | 3,004 ms |
| Cost for the run | $0.34, or $0.11 per thousand | $7.44, or $2.42 per thousand with caching |
Opus leads by 3.3 points, with a bootstrap interval of 2.3 to 4.4, and the two agree on 89.3% of utterances. Jev is 13 times faster and 22 times cheaper on the cached Opus price, and the study notes that a cascade, Jev first and Opus only on low-confidence cases, is an upper-bound result because it reused the same test set.
Source: OpenRouter — Is Jev as accurate as frontier models at classification? (Banking77, Jev 1.13 vs Claude Opus 5), September 24, 2026.
LiteLLM classified 80 authored prompts into four complexity tiers (simple, medium, complex, reasoning), three runs each, 240 calls per classifier, to pick a model route. The labels were not independently reviewed and downstream answer quality was not measured.
| Jev | Claude | |
|---|---|---|
| Matched the expected tier | 95.00% (228 of 240) | 73.75% (177 of 240) |
| Median latency | 126.81 ms | 688.40 ms |
| p95 latency | 231.16 ms | 896.94 ms |
| Cost for 240 calls | $0.0077 | $0.1985 |
On a routing task written by the people who run the router, Jev matched the intended tier far more often, 5.43 times faster at 96% lower cost. The authors say the result depends on their prompts and tier definitions, and the two classifiers agreed on only 78.75% of calls.
Source: LiteLLM — Jev classifier: 5.43x as fast as Haiku, 96% lower cost, September 18, 2026.
BERI classified 2,000 PhishNChips v5.2 emails, half phishing, with a 1,000-email holdout. Each model was asked once for a verdict, then asked five narrow questions whose answers were combined with a logistic regression fitted on the other half.
| Jev | Claude | |
|---|---|---|
| Asked once | 62.6% | 81.3% |
| Five questions, combined | 95.0% | 93.2% |
| Cost per 1,000 emails, one question | $0.038 | $0.462 |
| Cost per 1,000 emails, five questions | Included at the same rate | $1.02 |
Asked for a single verdict, Haiku is well ahead. Split into narrow questions, Jev gains 32 points and edges Haiku, though the 95.0 against 93.2 difference is not significant on the holdout (p = 0.063). The same study measured Jev’s calibration error at 0.107 on synthetic support tickets, 4.4 times the noise floor, with yes/no answers underconfident and choice and score answers overconfident.
Source: BERI — Jev scores 62.6% asked once and 95% split five ways, October 2026.
What the measurements add up to, and what they leave out.
The accuracy picture is consistent across the three studies. On a hard, flat label set like Banking77, a frontier Claude model is a few points better than Jev asked with one-line criteria: 84.4 against 81.0. On tasks whose labels were written by the team using them, or that decompose into narrow yes/no checks, Jev is level with or ahead of Haiku 4.5, and the phishing result shows how much the question design matters: 62.6% asked once, 95.0% asked five ways.
The speed and cost picture is not close. Every study measured Jev between 5 and 13 times faster at the median, with p95 latencies under 300 ms against 0.9 to 3 seconds for Claude, and between 12 and 26 times cheaper per request. The gap comes from the shape of the output: Jev reads the input once and emits a handful of probabilities, while Claude generates tokens, and Anthropic bills output at five times input. TypeSafe’s own claim of 193 times faster and 445 times cheaper on its four workflows has not been reproduced outside TypeSafe; the independent figures above are the ones to budget on.
What Jev cannot do is explain, and Arize puts the trade-off plainly: the biggest loss in moving a judge from an LLM to Jev is the explanation. Its recommendation is to run Jev on every trace for broad, comparable measurement, then send samples of the failures through Claude for the written reasoning. That cascade, Jev first and Claude on low confidence or for the explanation, is also what the OpenRouter study found to be the best accuracy per dollar.
List prices on October 7, 2026: Jev at $0.042 per million input tokens with output free; Claude Opus 5.5 at $4 input and $20 output per million, Sonnet 5.5 at $2 and $10, Haiku 4.5 at $1 and $5. A classification call that returns a 20-token JSON label on Claude pays for those output tokens at five times the input rate; on Jev it pays for the input only.
JevBench v1.6.0 measured the hosted Jev API at 0.24s median / 0.30s p95 per request, with a calibration score of 90.6; the Jev 1.13.0 guide has the version’s limits and the pricing page has Jev AI’s own plans.
Every study above says the same thing in its caveats: the result depends on the task and the prompt. Run your own routing, judging or screening cases on Jev here with no setup, keep the answers, and compare them with what Claude returns on the same inputs.
Not on hard, flat label sets asked with one-line criteria: on Banking77, Claude Opus 5 scored 84.4% against Jev 1.13’s 81.0%. On a routing task with author-written tiers Jev matched the intended tier 95.00% of the time against Claude Haiku 4.5’s 73.75%, and on phishing split into five narrow questions Jev reached 95.0% against Haiku’s 93.2%. Accuracy depends on the task and on how the question is written.
In the three independent studies on this page, 5.43 to 13 times faster at the median and 12 to 26 times cheaper per request. TypeSafe’s own claim of 193 times faster and 445 times cheaper has not been reproduced outside TypeSafe.
For the score, often; for the explanation, no. Jev returns a probability and no reasoning. Arize’s recommendation is to run Jev on every trace and send samples of the failures to an LLM for the written explanation.
Yes, and it is the pattern the studies point to: Jev decides which Claude model to route to, judges Claude’s outputs, and guards what a Claude agent reads; Claude writes the text. The LLM router and LLM-as-a-judge use cases on this site are ready judges for both halves.
No. Jev accepts 64k tokens per request, with 32k for the state; current Claude models accept 1M. Filter long inputs before sending them to Jev; TypeSafe says accuracy falls when the state is full of unrelated text.
No. Jev is developed by TypeSafe. Claude is Anthropic’s. Jev AI, this site, is an independent playground and API for Jev and is affiliated with neither company.
Every figure on these pages was read from its source on October 7, 2026.
Claude and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Jev is developed by TypeSafe; Jev AI is an independent playground and API. Figures are quoted from the sources above as they read on October 7, 2026 and can change.