Jev AI

Jev-as-a-Judge · guide with evidence · published October 7, 2026

Jev-as-a-Judge

Grading with typed verdicts and calibrated probabilities instead of written critiques: what has been measured, what we measured ourselves, and how to set one up.

An LLM judge reads an answer and writes an opinion. A Jev judge reads the same answer and returns a probability for each question you asked: supported or not, which failure category, how complete. That makes the verdict cheap, fast and thresholdable, and it makes the explanation disappear. This page collects the independent measurements that exist, adds one of our own with the data and script published, and ends with a setup that uses the probability properly: accept when confident, escalate when unsure.

91.3%accuracy on our 300 HaluEval verdicts
$0.021per thousand verdicts in that run
91.5%agreement with Claude Fable 5.1 over 6,003 rubric checks (Good Start Labs)
0.36%of GPT-6’s judging fee (Carnegie Mellon)

What Jev-as-a-Judge means

The same job as LLM-as-a-judge, with a different kind of answer.

You still supply the question, the reference and the candidate answer. The difference is the question you ask the judge. An LLM is asked to rate and explain; Jev is asked typed questions and returns, for each, a probability over the options you defined. A Noul gives the probability that a statement holds. A Choice picks a verdict category and shows every category’s probability. A Score places the answer on ordered rubric levels you wrote, as a probability-weighted value.

The Carnegie Mellon paper that named the pattern adds the operating rule: use the probability. Accept the verdict when the judge is confident, escalate to a reasoning judge or a person when it is not. Everything below is about how well that works and where it does not.

What has been measured

Five measurements by five different parties, quoted with their setup and their caveats. None used the same data, so read them as five views of one model rather than one ranking.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Yubo Li, Yidi Miao, Ramayya Krishnan and Rema Padman, Carnegie Mellon University · arXiv, revised October 6, 2026

Setup
A decision-only judge against sixteen generative and reward-model judges, with a confidence threshold that accepts its verdict or escalates to a reasoning judge
Accuracy
Within three points of GPT-6 where the verdict can be read from the text; further behind on math, code and logic, where it must be derived
Cascade
0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee; a pre-specified live test matched GPT-6 exactly
Cost and speed
0.36% of GPT-6’s fee as a judge; 0.15-second median latency

The paper is the strongest independent evidence, and it is explicit about the failure mode: confidence routing weakens on style-adversarial pairs and reference-free prose.

Read the source ↗

6,003 rubric checks against five LLM judges

Good Start Labs, reported by Langfuse · September 18, 2026

Setup
The same rubric instructions graded by Jev and five LLMs over 6,003 checks
Agreement with Claude Fable 5.1
Jev 91.5%; DeepSeek V4.1 Flash 93.5%
Cost per million graded answers
Jev $160; DeepSeek V4.1 Flash $260; GPT-5.6 Luna $400; Gemini 3.8 Flash $1,600; Claude Fable 5.1 $33,000

Agreement with a frontier judge is not accuracy against humans, and the cheapest LLM agreed slightly more often. The cost gap is the finding.

Read the source ↗

30 human-labelled QA answers in MLflow

Yuki Watanabe, Databricks, on the MLflow blog · September 22, 2026

Setup
One Noul question: is the candidate answer factually and technically correct for the question, using the supplied MLflow documentation
Agreement with humans
Jev 1.13.0 30 of 30; GPT-5.6 Terra 30 of 30; DeepSeek-V4.1-Flash 29 of 30; Claude Sonnet 4.6 27 of 30
Median latency
Jev 369 ms; DeepSeek 910 ms; GPT-5.6 Terra 1,091 ms; Sonnet 4.6 1,610 ms
Cost per 1,000
Jev $0.0247; DeepSeek $0.0624; GPT-5.6 Terra $0.896; Sonnet 4.6 $1.67

Thirty examples is a smoke test, not a benchmark. The author’s own conclusion is that Jev suits production monitoring more than development iteration, because it returns no rationale.

Read the source ↗

Can decision models replace LLM judges?

Arize, quoting TypeSafe’s four-workflow evaluation · September 2026

Setup
TypeSafe’s own evaluation across four workflows, scored by agreement with a frontier reference
Agreement
Jev 68% at $0.0004 and 0.4 s per case; GPT-5.6 Terra 68% at $0.03 and 10 s; Claude Opus 5 73% at $0.18 and 38 s

These are the vendor’s numbers. Arize’s advice is the pattern this page recommends: run Jev on every trace, then send samples of the failures to an LLM for the written explanation.

Read the source ↗

One question or five

BERI · October 2026

Setup
2,000 PhishNChips emails; each judged once with a single verdict, then with five narrow questions combined by a fitted logistic regression
Accuracy
Asked once 62.6% (Claude Haiku 4.5: 81.3%); five questions combined 95.0% (Haiku: 93.2%)
Calibration
Expected calibration error 0.107 on synthetic support tickets, 4.4 times the noise floor; yes/no answers underconfident, choice and score answers overconfident

The largest effect anyone has measured on a Jev judge is question design, not model choice. It also shows why thresholds need checking on your own labels.

Read the source ↗

Langfuse evaluators and Braintrust scorers both offer Jev as a judge model; Braintrust quotes TypeSafe’s own figure of up to 444.6 times lower cost, which no third party has reproduced, and notes that Jev’s confidence values are not operating thresholds out of the box.

Our own test: 300 verdicts on HaluEval

A first-hand measurement, run on 2026-10-07, with the data, the question and the script published so it can be repeated.

DatasetHaluEval QA: 150 items, every 66th of 10000 items; each item has a knowledge passage, a question, the correct answer and a hallucinated answer, so every verdict has a label
Verdicts300: one Noul question per answer, “Is the candidate answer correct for the question, according only to the knowledge passage?”
Modeljev-latest (jev-1.13.0) through TypeSafe’s API at list price, concurrency 4, 2026-10-07
Accuracy at a 0.5 threshold91.3%
Correct answers accepted95.3%
Hallucinated answers rejected87.3%
Latency535 ms median, 1088 ms p95; Client-observed round trip from a laptop in Asia at concurrency 4, including the network hop.
Tokens and cost500 input tokens per verdict; $0.0063 for the run, $0.021 per thousand verdicts at $0.042 per million
CalibrationExpected calibration error 0.056 over ten probability bins

Calibration, bin by bin

For each band of the probability Jev returned, how often the answer really was supported. A calibrated judge sits close to the diagonal.

Probability returnedVerdictsMean probabilityShare actually supported
0.0–0.11100.0240.045
0.1–0.2100.1550.000
0.2–0.3120.2420.167
0.3–0.430.3670.000
0.4–0.530.4400.000
0.5–0.650.5420.800
0.6–0.750.6320.600
0.7–0.8110.7470.727
0.8–0.9100.8500.600
0.9–1.01310.9730.931

What we found

  • Jev accepted 95.3% of correct answers and rejected 87.3% of hallucinated ones from one question and a two-line rubric, with no examples and no fine-tuning.
  • The two outer bins hold 241 of 300 verdicts and are close to calibrated. The 0.8–0.9 band is not: a mean probability of 0.85 against 0.60 actually supported, which is where an accept threshold should not sit.
  • The misses are mostly hallucinated answers that stay vague or assert something the passage never mentions, rather than answers that flatly contradict it. A second question aimed at exactly that, “does the answer add a fact the passage does not contain?”, is the obvious next decomposition.
  • Latency from a laptop in Asia was 535 ms at the median, slower than the 0.15 to 0.37 seconds measured closer to the API; the network hop, not the model, is most of it.

Three of the misses

  • Hallucinated answer accepted · p = 0.93When was the company that publishes the journal Magnetic Resonance Imaging established ?The company that publishes Magnetic Resonance Imaging was established during the 19th century.
  • Correct answer rejected · p = 0.24What Bible college located in Louisville, Kentucky, was attended by Paul R. House who served for a time as president of the Evangelical Theological Society?Southern Baptist Theological Seminary
  • Hallucinated answer accepted · p = 0.92Are Onew and Judith Durham both singers?While Onew is a singer and actor, Judith Durham is both a singer and musician.

Limits of this test. One dataset, one domain, one question wording, one run. HaluEval’s hallucinated answers were generated by a model and are often blatant, so 91.3% is an upper bound for subtler production errors. No human relabelling was done; the labels are the dataset’s own. Latency includes a long network path. The script, the sampling rule and the full bin table are in the repository file named above.

How to set up a Jev judge

One request, three typed questions, and a threshold rule. The LLM-as-a-judge template opens this in the playground; “Get API code” exports it.

curl --fail-with-body https://jev-ai.pro/api/v1/systemone \
  -H "Authorization: Bearer $JEV_AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": {
      "question": "What is the return window for unopened items?",
      "reference": "Unopened items may be returned within 30 days of delivery for a full refund.",
      "candidate_answer": "You can return unopened items within 30 days, and we also cover return shipping."
    },
    "questions": {
      "support": { "type": "choice", "instructions": "Does the reference support every factual claim in the answer?",
                   "criteria": { "supported": "Every claim is stated in the reference.",
                                 "contradicted": "A claim conflicts with the reference.",
                                 "unsupported": "A claim is neither stated nor denied by the reference." } },
      "relevant": { "type": "noul", "instructions": "Does the answer address the question that was asked?" },
      "complete": { "type": "score", "instructions": "How much of what the question asks for does the answer cover?",
                    "criteria": ["Misses the main point", "Covers the main point, misses a required detail", "Covers everything asked"] }
    }
  }'
  1. 01

    Separate the criteria

    Source support, relevance and completeness are three questions, not one rating. BERI’s 62.6% to 95.0% jump came from this alone. Each question is evaluated independently against the same state; combine them in code.

  2. 02

    Give the judge a way out

    A Noul cannot abstain. Where the reference may not settle a claim, use a Choice with an explicit unsupported or cannot-tell option, as above, so the judge is never forced to call a missing fact true or false.

  3. 03

    Fit the thresholds on your own labels

    Grade a few hundred human-labelled examples, plot accuracy against the returned probability, and set two lines: accept above one, reject below the other, escalate between. Our run says not to put the accept line at 0.8; yours may differ.

  4. 04

    Cascade for the explanation

    Jev grades every trace; low-confidence and failing cases go to an LLM for the written reason, or to a person. Carnegie Mellon measured that cascade at GPT-6’s accuracy for 41% of its fee.

  5. 05

    Pin and batch

    Save the questions as a judge and name jev-1.13.0 so a future release does not move your scores. The batch page runs one judge over a CSV or JSONL of answers.

Typed judge, LLM judge or a code check

Three graders that belong in the same evaluation stack.

 Jev judgeLLM judgeCode check
OutputTyped verdict with a probabilityScore with a written reasonPass or fail
Variance between runsLowHighNone
Cost and latencyLow: 0.2 to 0.5 s, cents per thousandHighest: seconds, dollars per thousandLowest
Best forHigh-volume gates, monitoring every trace, routing to reviewNuanced quality, tone, anything a person will readExact formats, schemas, totals, dates
Weak atDerived verdicts (math, code, logic); no rationaleCost at volume; drift between runsAnything that is not literal

Both typed and generative judges can be moved by instructions hidden in the text they grade. Screen untrusted content with the prompt-injection judge first, and keep arithmetic and schema checks in code.

Where Jev-as-a-Judge fails

From the studies above and our own misses.

  • Derived verdicts. Proofs, code correctness, multi-step arithmetic: the answer has to be worked out, and Jev reads rather than reasons. Carnegie Mellon found it furthest behind GPT-6 exactly there.
  • No explanation. A failing score tells you nothing about why. Every integration guide says the same: keep an LLM in the loop for the reason.
  • Cannot abstain on yes/no. A Noul always answers. Use a Choice with a cannot-tell option where the evidence may be missing.
  • Context rot. Accuracy falls when the state is full of unrelated text. Send the passages that matter, not the whole document.
  • Calibration is uneven. Good at the extremes, overconfident in the upper-middle bands in our run and in BERI’s; fit thresholds on your data.
  • Style and reference-free prose. The paper found confidence routing weakens on style-adversarial pairs and prose with no reference to check against.

Jev-as-a-Judge FAQ

What is Jev-as-a-Judge?

Using TypeSafe’s Jev, a decision model, as the grader in an evaluation loop. Instead of asking an LLM to write a critique, you give Jev the question, the reference and the candidate answer and ask typed questions: is it supported, which failure category, how complete. Each answer comes back as a probability, so a threshold decides what passes, what fails and what a person or a reasoning model should look at.

How accurate is Jev as a judge?

It depends on whether the verdict can be read from the text. Carnegie Mellon found Jev within three points of GPT-6 on such tasks and further behind on math, code and logic. MLflow saw 30 of 30 agreement with human labels on documentation QA. Our own test on 300 HaluEval verdicts scored 91.3%, rejecting 87.3% of hallucinated answers. Decomposing the judgment into narrow questions raised BERI’s accuracy from 62.6% to 95.0%.

How much does a Jev judge cost compared with an LLM judge?

Jev bills input tokens only, at $0.042 per million. Good Start Labs measured $160 per million graded answers against $33,000 for Claude Fable 5.1; MLflow measured $0.0247 per thousand against $0.896 for GPT-5.6 Terra; our run cost $0.021 per thousand verdicts at about 500 tokens each.

Does Jev explain its verdict?

No. It returns probabilities and nothing else. Every study on this page names that as the main loss, and the fix is the cascade: let Jev grade every trace, then send the failures, or a sample of them, to an LLM for the written reason.

Can a Jev judge say it cannot tell?

A Noul (yes/no) question cannot abstain; it always returns a probability. Use a Choice question with an explicit insufficient-evidence option when the reference may not settle the question, and treat mid-range probabilities as a signal to escalate rather than a verdict.

Are the probabilities calibrated enough to use as thresholds?

Roughly, and not uniformly. JevBench v1.6.1 scores Jev’s calibration at 90.6, our run measured an expected calibration error of 0.056, and BERI measured 0.107 on harder tickets with overconfidence on choice and score answers. Fit the accept and escalate thresholds on a few hundred of your own labelled examples; the Carnegie Mellon paper gives a recipe.

Which tasks should not use Jev as a judge?

Anything where the verdict has to be derived rather than read: checking a proof, running code in your head, multi-step arithmetic. Also reference-free prose quality and style, where the paper found confidence routing weakens, and any audited or customer-facing verdict that needs a written reason.

Where can I run a Jev judge?

In the Jev AI playground and API on this site, or through TypeSafe’s API, OpenRouter, Vercel AI Gateway or Cloudflare Workers AI. Langfuse evaluators and Braintrust scorers both offer Jev as a judge model; the LLM-as-a-judge use case on this site is a ready template you can edit and export.

About this page

How it was made, so you can judge it.

Who. Jev AI operates an independent playground and API for Jev and is not affiliated with TypeSafe, Carnegie Mellon, Langfuse, MLflow, Braintrust, Arize or BERI. We sell access to the model this page is about; the third-party figures are quoted so that claim is not ours alone, and our own test is published in full so it can be checked.

How. Third-party figures were read from the linked sources on October 7, 2026. Our measurement used TypeSafe’s public API at list price, the public HaluEval QA split under its MIT licence, and the script in the repository; nothing was filtered after the run.

When. Published October 7, 2026. Jev is versioned, and jev-latest currently resolves to jev-1.13.0; a new release would make the numbers above historical, and the page will say so when that happens.

Sources

Jev is developed by TypeSafe. Claude, GPT, DeepSeek, Gemini and the products named above belong to their respective owners and are not affiliated with Jev AI.