On this site
Jev 1.13.0
TypeSafe · System One decision model
Any question, any labels, no training data; 64k tokens of context; a calibrated probability per option.
Jev vs a generative model · independent measurements · updated October 7, 2026
The same job as a fine-tuned classifier, with the labels written at request time instead of trained in.
A BERT classifier and Jev produce the same thing: scores over a set of labels for a piece of text. The difference is where the labels come from. BERT learns them from your labelled data in a training run and answers nothing else; Jev reads the question, the options and the rubric in the request and answers whatever you ask. That trade, agility against a trained fit, is the whole comparison. This page quotes the measurements that exist, including where a fine-tune is still ahead, and the one BERT-family decision model that takes Jev’s request format.
On this site
TypeSafe · System One decision model
Any question, any labels, no training data; 64k tokens of context; a calibrated probability per option.
vs
Compared with
Google and the open-source community · Fine-tuned encoder classifiers: BERT-base, DeBERTa, ModernBERT
A fixed label set learned from your data; milliseconds on your own GPU; the best accuracy on a stable, narrow task.
Jev and BERT share no benchmark, so every number here is one party’s own test, quoted with its setup and its caveats.
OpenRouter ran Jev 1.13 on the 3,080-utterance Banking77 test split with one-line criteria per intent and no fine-tuning. Fine-tuned encoder results on the same benchmark, trained on its full labelled training set, are published in the low 90s.
| Jev | BERT | |
|---|---|---|
| Accuracy | 81.0% (95% interval 79.6 to 82.3) | Low 90s, fine-tuned on about 10,000 labelled utterances |
| Macro-F1 | 80.5% | Not reported in the same study |
| Training data used | None; 77 one-line criteria | The full Banking77 training split |
| Median latency | 175 ms over the network | Milliseconds on a GPU |
With a stable label set and ten thousand labelled examples, a fine-tuned encoder is about ten points ahead. That is the cost of one-line criteria instead of a training run, and the case for a fine-tune on a narrow, settled task.
Source: OpenRouter — Is Jev as accurate as frontier models at classification? (Banking77, Jev 1.13 vs Claude Opus 5), September 24, 2026.
BERI pre-registered 50 predictions before collecting data, then ran 5,721 calls against a 400-item test set, comparing Jev zero-shot with hand-written keyword rules and a supervised TF-IDF classifier.
| Jev | BERT | |
|---|---|---|
| Accuracy | 95.9% zero-shot | 66.0% supervised TF-IDF; 77.2% hand-written keywords |
| With wrong criteria | 16.7%, below the 25% random floor | Not applicable |
| Calibration error, synthetic tickets | ECE 0.107, 4.4x the 0.024 noise floor | Softmax scores usually need temperature scaling |
Against a classical supervised baseline, zero-shot Jev wins by 30 points. The same study shows the failure mode a trained classifier does not have: criteria written wrongly take Jev below chance, because the request is the model’s only definition of the task.
Source: BERI — Jev scores 62.6% asked once and 95% split five ways, October 2026.
XenoSpectrum collects published results on fine-tuned encoders against prompted generative models; none of them include Jev, and the article says no third-party benchmark of Jev against fine-tuned encoders under identical labels exists yet.
| Jev | BERT | |
|---|---|---|
| Legal contract clauses | Not measured | Fine-tuned DeBERTa 87.8% against zero-shot GPT-4 67.2% |
| 8-class political manifestos | Not measured | BERT fine-tuned on 200 samples 53.9% against prompting 48.8% |
| 20-class COVID policy | Not measured | Fine-tuned on 500 samples 65.7% against prompting 65.8% |
| Phishing, five narrow questions | 95.0% (Jev), 93.2% (Claude Haiku 4.5) | Not measured |
On specialised domains with a few hundred labelled examples, a fine-tuned encoder has repeatedly beaten prompted generative models. Jev has not been put on those tasks. Its own decomposition result, 95.0% on phishing from five narrow questions, is the number to test against a fine-tune on your data.
Source: XenoSpectrum — Jev, BERT classifiers and decomposition, September 2026.
What the measurements add up to, and what they leave out.
BERT-base is a 110-million-parameter encoder with a 512-token window; bert-base-uncased was still downloaded 47 million times a month in September 2026, and ModernBERT extended the idea to 8k tokens. A classifier built on it is a head trained on your labelled examples. Once trained it is fast, cheap, private and, on the task it was trained for, usually the most accurate thing you can run: the Banking77 gap, low 90s against Jev’s 81.0%, is typical of a stable task with ten thousand labels.
Jev inverts the order. There is no training run: the question, the options and the rubric arrive in the request, so the first decision costs nothing to set up and the label set can change tomorrow. Against classical supervised baselines on a pre-registered test it scored 95.9% to their 66.0%. The price is that the request is the only definition of the task: wrong criteria took it to 16.7%, below chance, and BERI measured a calibration error of 0.107 on synthetic tickets, so thresholds still want checking on your own labels.
The honest comparison is sequential rather than either-or. Start on Jev, where no labelled data exists yet, and keep the answers. If the task settles into something narrow and high-volume, those stored decisions are the labelled set that trains an encoder, and Laya, Convai’s ModernBERT-large decision head that takes Jev’s request shape, is the BERT-family model built for exactly that hand-off: 33 to 40 ms on a T4 once fine-tuned, and below a majority-class baseline before.
Jev lists at $0.042 per million input tokens with output free, $0.11 per thousand Banking77 requests in the OpenRouter study. A fine-tuned encoder has no per-request price: you pay for the labelling, the training run and the GPU or CPU that serves it, and the marginal decision is close to free.
JevBench v1.6.0 measured the hosted Jev API at 0.24s median / 0.30s p95 per request, with a calibration score of 90.6; the Jev 1.13.0 guide has the version’s limits and the pricing page has Jev AI’s own plans.
Every study above says the same thing in its caveats: the result depends on the task and the prompt. Run your own routing, judging or screening cases on Jev here with no setup, keep the answers, and compare them with what BERT returns on the same inputs.
A BERT classifier learns a fixed label set from your labelled data and must be retrained when the labels change. Jev takes the question, options and rubric at request time, so a new rule needs no training run. A fine-tuned encoder is still faster, runs on your own hardware and, on a stable narrow task, usually more accurate.
Not when BERT has been fine-tuned on the task: on Banking77, fine-tuned encoders score in the low 90s against Jev’s 81.0% zero-shot. Against classical supervised baselines without a fine-tune, Jev scored 95.9% to a TF-IDF classifier’s 66.0% on a pre-registered test. No third-party benchmark has yet put Jev and a fine-tuned encoder on identical tasks and labels.
No. TypeSafe does not disclose Jev’s architecture or size; it describes a post-trained language model that reads a state and answers typed questions in parallel. Laya is the BERT-family model that takes Jev’s request shape: a decision head on ModernBERT-large.
No. It answers zero-shot from the criteria in the request. The flip side is that wrong criteria are catastrophic: BERI measured 16.7% accuracy, below chance, with a mis-written task definition.
BERT on a GPU, by an order of magnitude: milliseconds, or 33 to 40 ms for Laya on a Tesla T4, against Jev’s 175 ms median over the network in the OpenRouter study. Jev is fast for a hosted model; a local encoder is faster still.
Jev is post-trained for calibration and scores 90.6 on JevBench v1.6.0’s calibration axis, but BERI measured an expected calibration error of 0.107 on synthetic tickets, with yes/no answers underconfident and choice and score answers overconfident. A BERT softmax usually needs temperature scaling on held-out data. Check thresholds on your own labels either way.
Every figure on these pages was read from its source on October 7, 2026.
BERT and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Jev is developed by TypeSafe; Jev AI is an independent playground and API. Figures are quoted from the sources above as they read on October 7, 2026 and can change.