Model guide · open Jev alternative · published October 9, 2026
Kev
Jev’s request format on open Qwen weights, in four sizes, with a fine-tuning path Jev does not have.
Kev is the most-starred open reconstruction of Jev: a small adapter and pointer head on Qwen3.5 and Qwen3.8 that answers typed questions with calibrated probabilities and never generates text. TypeSafe’s own SDK works against a Kev server unchanged. This page sets out what it is, the author’s measurements against Jev, where its earlier checkpoints sit on JevBench, how to run and fine-tune it, and what we measured ourselves by running Kev-4B on a laptop against hosted Jev on the same 300 verdicts.
What Kev is
An architecture, four checkpoints and a server, not a hosted service.
Each Kev checkpoint is a rank-16 LoRA adapter plus a small pointer head on a Qwen base. The state and a question are laid out as one token sequence with the options marked; the head scores every option’s closing marker against the question’s final decision token, and a softmax turns the scores into probabilities. Nothing is decoded. Because Qwen3.5 and 3.8 mix attention with recurrent Gated DeltaNet layers, each question runs as its own row over the shared state rather than in one masked pass, and the author reports that asking questions together or separately changes probabilities by less than 4e-6 in fp32.
The 0.8B, 4B and 9B start from Qwen base checkpoints and share one recipe: public decision datasets, generated policy and rule examples, a document stage on 5,219 real consumer-finance complaints, a skills stage, then one fitted temperature. Kev-27B starts from Qwen’s post-trained 3.8-27B release, whose post-training data is unknown, and trains every backbone weight. The project describes itself as a reconstruction of the architecture in “Jev’s Architecture Unmasked” and is not affiliated with TypeSafe.
The four sizes
Released together as Kev 1.0 with a v1.0 tag on each Hub repository. The index is the author’s held-out composite; Jev scores 54.0 on it.
| Model | Base | Runs on: CUDA | Runs on: Mac (MLX) | Validated context | Held-out index |
|---|---|---|---|---|---|
| Kev-0.8B | Qwen3.5-0.8B-Base | L4 or any 4 GB GPU | Any Apple-silicon Mac | 8,192 | 23.3 |
| Kev-4B | Qwen3.5-4B-Base | L40S, H100 | 32 GB Mac | 8,192 | 38.0 |
| Kev-9B | Qwen3.5-9B-Base | L40S, H100 | 32 GB Mac or larger (expected) | 8,192 | 41.0 |
| Kev-27B | Qwen3.8-27B, post-trained | B200, H200, H100 80 GB | 96–128 GB Mac (expected) | 65,536 | 52.3 |
The author’s guidance: start with Kev-4B; move to 9B with a bigger GPU, to 27B with an 80 GB card for the most accurate Kev, and to 0.8B when size matters more than accuracy.
Accuracy by the author’s own suites
“New sources” are datasets and rule types Kev never trained on, the closest thing to your own questions. Cells are development / test; lower Brier is better.
| Model | New sources, dev | New sources, test | Trained sources | Brier, new sources |
|---|---|---|---|---|
| Kev-0.8B | 0.648 | 0.697 | 0.827 / 0.838 | 0.481 / 0.416 |
| Kev-4B | 0.817 | 0.838 | 0.873 / 0.865 | 0.269 / 0.242 |
| Kev-9B | 0.820 | 0.852 | 0.874 / 0.873 | 0.289 / 0.217 |
| Kev-27B | 0.851 | 0.889 | 0.865 / 0.866 | 0.225 / 0.156 |
| Jev (hosted) | 0.857 | – | 0.845 / – | 0.211 / – |
The author adds that Kev-27B is within three points of Jev, or ahead, on 9 of 11 new-source categories; that knowledge questions track the base model, with Kev-9B at 0.73 on MMLU against Jev’s 0.90 and Kev-27B at 0.565 on MMLU-Pro against 0.840; and that at a 5% error budget Kev-4B to 27B can automate 0.52 to 0.69 of new-source decisions against Jev’s 0.70. These are the project’s measurements on its own suites, with Jev called over the API.
Kev on JevBench v1.6.1
Benchmark Heaven measured four earlier Kev checkpoints under one method. None is a Kev 1.0 size.
| System | Rank of 127 | Score | Intelligence | Calibration | Speed |
|---|---|---|---|---|---|
| Jev 1.13.0 (TypeSafe) | #3 | 71.5 | 63.6 | 90.6 | 91.5 |
| kev 8B (research preview) | #50 | 10.6 | 29.3 | 70.8 | 89.0 |
| kev 4B (research preview) | #57 | 6.6 | 19.5 | 68.5 | 90.5 |
| kev 0.6B (research preview) | #66 | 0.5 | 7.6 | 66.3 | 91.5 |
| kev 0.5B (research preview) | #78 | 0.1 | 4.6 | 61.2 | 92.4 |
Those rows are research previews named 0.5B, 0.6B, 4B and 8B, measured before Kev 1.0 shipped its 0.8B, 4B, 9B and 27B on 24 September; the benchmark has not re-measured the release. Until it does, the author’s suites above and our test below are the only numbers for Kev 1.0, and neither is comparable with a JevBench score. The full board is on the JevBench results page.
Our own test: Kev-4B on a laptop against hosted Jev
The 300 verdicts from our Jev-as-a-Judge test, replayed on Kev-4B running locally. Same items, same question, same threshold.
| Dataset and question | The same 300 verdicts as our Jev-as-a-Judge test: 150 HaluEval QA items, each with a correct and a hallucinated answer, one Noul question, “Is the candidate answer correct for the question, according only to the knowledge passage?” |
|---|---|
| Kev side | Kev-4B, local MLX bf16 on an Apple-silicon Mac with 48 GB, served by the project’s own kev.serve at its shipped temperature, 2026-10-08 |
| Jev side | jev-latest (jev-1.13.0) through TypeSafe’s API at list price, 2026-10-07 |
| Kev-4B, local | Jev 1.13.0, hosted | |
|---|---|---|
| Accuracy at a 0.5 threshold | 81.7% | 91.3% |
| Correct answers accepted | 96.0% | 95.3% |
| Hallucinated answers rejected | 67.3% | 87.3% |
| Expected calibration error | 0.125 | 0.056 |
| Median latency per verdict | 559 ms on the laptop | 535 ms over the network from Asia |
| p95 latency | 636 ms | 1088 ms |
| Input tokens per verdict | 214 | 500 |
| Cost of the run | Electricity | $0.0063 |
What we found
- On this reference-checking task Kev-4B lands 9.6 points below hosted Jev: 81.7% against 91.3%, accepting 96.0% of correct answers and rejecting 67.3% of hallucinated ones, from a 4B model on a laptop.
- Calibration: Kev’s expected calibration error was 0.125 against Jev’s 0.056. Verdicts Kev put in the 0.9–1.0 band were supported 94% of the time; those in the 0.0–0.1 band 5%.
- Where the gap sits: Kev-4B accepted 32.7% of the hallucinated answers, Jev 12.7%. Kev put 47 verdicts in the 0.5 to 0.8 bands with a mean probability of 0.64, of which only 23% were actually supported; Jev put 21 there, 71% supported. A threshold above 0.8 would fix most of Kev’s misses at the cost of recall, which is exactly what fine-tuning on your own labels is for.
- Latency is a different kind of number: 559 ms per verdict is model time on an Apple-silicon laptop through MLX, while Jev’s 535 ms includes a network hop from Asia. The author measures Kev-4B at 18.1 ms on an H100.
- Each verdict cost Jev about 500 input tokens, $0.021 per thousand; Kev’s cost is the machine it runs on.
Limits of this test
- One dataset of short question-answer pairs with blatant, model-generated hallucinations; 81.7% and 91.3% are both upper bounds for subtler production errors.
- One question wording, written for Jev and reused verbatim on Kev; neither was tuned.
- One run, bf16 on MLX, at the shipped temperature; the author’s evaluation path is fp32 on a GPU.
- Labels are the dataset’s own; nothing was relabelled or filtered after the run. The script and both result files are in the repository.
Run Kev locally
Python 3.12 or 3.13 and uv. The server speaks the same request shape as Jev, so a Jev AI request body works with the model name changed.
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009
# CUDA or ROCm on a GPU, MLX on Apple silicon; the first run downloads the adapter and the base.
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"model": "kev-latest",
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"}
}}'curl --fail-with-body https://jev-ai.pro/api/v1/systemone \
-H "Authorization: Bearer $JEV_AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "jev-latest", "state": "…", "questions": {"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"}}}'
# Same body against a Kev server: drop the key, set "model": "kev-latest", change the host.To host it, the kev-deploy skill serves Kev-4B on a Modal L40S behind a bearer key and scales to zero when idle. The running Jev locally guide puts Kev beside the other open alternatives by hardware.
Fine-tune Kev on your own workload
The thing Kev can do that Jev cannot: learn your labels and recalibrate on them.
# With a coding agent: install the skill, then ask it to "fine-tune Kev on my support tickets"
npx skills add jaredpalmer/kev@kev-finetune
# By hand: train.jsonl holds your labelled decisions (state, question, options, target)
uv run python -m kev.train --data train.jsonl --base Qwen/Qwen3.5-4B-Base --init_from jaredpalmer/kev-4b --out runs/mine
uv run python -m kev.benchmark --run runs/mine --data heldout.jsonl --out runs/mine-eval
uv run --extra serve python -m kev.serve --run runs/mine --port 8009- 01
Interview and extract
The kev-finetune skill scans a codebase for Jev and TypeSafe calls, lifts the question literals and finds labelled CSV or JSONL files, then asks what the decision is and what error rate you can tolerate.
- 02
Data
Convert existing labels, or let an OpenAI-compatible model write balanced records from the spec; a size planner says how many records make the gain measurable. Rows split by state into train, calibration and development.
- 03
Train, calibrate, score
One Modal H100, about 12 minutes and $1 for 400 records with Kev-4B; a temperature is fitted on the calibration split and the result is scored against the released checkpoint. The README reports 67.7% to 73.6% on an example workload and 0.804 to 0.904 after one epoch on 5,219 real complaints.
- 04
Deploy and tear down
The trained run serves as a TypeSafe-compatible endpoint on Modal, scaling to zero after five idle minutes; the skill removes everything it created when you are done.
Limits to plan around
From the model cards and README.
- Length. The 0.8B, 4B and 9B trained mostly on states up to 384 tokens and are validated to 8,192; the server refuses anything over 65,536 with a 422 rather than truncating. Kev-27B is validated to 65,536.
- Knowledge. Questions whose answer depends on general knowledge track the base model: MMLU-Pro 0.565 for Kev-27B against Jev’s 0.840.
- Option order. Changing the order of options can change an answer; the /v1/systemone/permute route averages over orders at extra cost.
- English only, no day-precision date arithmetic without the KEV_DATE_FACTS preprocessor, and one global temperature rather than per-question calibration.
- Hardware. Kev-27B needs an 80 GB GPU or a 96 to 128 GB Mac, and its 27B base was post-trained on data the author cannot describe.
Kev or Jev?
Pick Jev if
- You want a hosted, versioned model with nothing to run, on day one, with no labelled data.
- Inputs run past 8k tokens and you are not on an 80 GB GPU.
- The task leans on general knowledge, or on the hardest reasoning tiers where JevBench’s intelligence axis separates systems.
Pick Kev if
- You have, or can make, a few hundred labelled examples of your exact questions and want probabilities calibrated on them.
- Your data must stay inside your network, or your volume makes a GPU you already own cheaper than tokens.
- You want weights you can inspect, pin and retrain, under Apache-2.0.
Use both when
- You start on Jev with no training set, keep every decision, and fine-tune Kev on those stored runs once the task settles.
- Kev serves the high-volume narrow questions locally and Jev takes the long, knowledge-heavy or rare ones.
Kev FAQ
What is Kev?
An open family of Jev-style decision models by Jared Palmer: a rank-16 LoRA adapter and a pointer head on a Qwen3.5 or Qwen3.8 base that answers noul, choice and score questions with calibrated probabilities and generates no text. Four sizes from 0.8B to 27B are released together as Kev 1.0 under Apache-2.0, and its server speaks TypeSafe’s /v1/systemone wire format.
Is Kev as accurate as Jev?
Close at the top size, by the author’s numbers: on datasets Kev never trained on, Kev-27B scores 0.851 against Jev’s 0.857, and Kev-4B 0.817. In our own 300-verdict HaluEval test, Kev-4B on a Mac scored 81.7% against hosted Jev’s 91.3%. JevBench v1.6.0 has only measured four earlier research-preview checkpoints, which rank #50 and below; Kev 1.0 is unmeasured there.
Can I use the TypeSafe SDK with Kev?
Yes. Point the TypeSafe Python client at a Kev server with any api_key and model kev-latest and keep the rest of the code; the request and response shapes are the same, and Kev adds a /v1/systemone/permute route that runs one choice question over several option orders.
What hardware does Kev need?
Kev-0.8B runs on any 4 GB GPU or any Apple-silicon Mac; Kev-4B on an L40S or H100, or a 32 GB Mac through MLX; Kev-9B on the same GPUs; Kev-27B on an 80 GB GPU such as an H100, H200 or B200, or a 96 to 128 GB Mac. Resident GPU memory for Kev-4B is 14.3 GB.
How do I fine-tune Kev?
With the kev-finetune skill for a coding agent, or by hand with python -m kev.train. The skill interviews you, extracts the questions your code already sends to Jev, converts or generates labelled records, trains and calibrates on one Modal H100 for about $1 and 12 minutes per 400 records, scores the result against the released checkpoint, deploys a TypeSafe-compatible endpoint and tears everything down. The README reports Kev-4B going from 67.7% to 73.6% on an example workload and from 0.804 to 0.904 after one epoch on 5,219 real complaints.
Why fine-tune Kev instead of calling Jev?
Because Jev is a fixed hosted model: on your own labels it is out of distribution and its probabilities cannot be recalibrated. Kev can be trained on a few hundred to a few thousand of your exact questions and its temperature fitted on a held-out slice, so the probabilities it serves are calibrated for your data and the weights stay inside your network.
How long a document can Kev read?
The server accepts states of up to 65,536 tokens, but the 0.8B, 4B and 9B models trained mostly on states of up to 384 tokens and are validated to 8,192; beyond that the author found accuracy falls past tolerance. Kev-27B trained on states of up to 32,768 and is validated to 65,536. Jev accepts 64k per request.
Is Kev affiliated with TypeSafe?
No. It is an independent Apache-2.0 reconstruction of the architecture described in “Jev’s Architecture Unmasked”, trained on public datasets and generated policy examples. Jev AI, this site, hosts Jev and is affiliated with neither project.
About this page
How it was made, so you can judge it.
Who. Jev AI operates an independent playground and API for Jev and sells access to it, which is the opposite side of the trade from a model you run yourself; that is why Kev’s own numbers are quoted with their sources and our test is published in full. We are not affiliated with Jared Palmer, TypeSafe or Benchmark Heaven.
How. Project facts are from the Kev repository, README and model cards as read on October 9, 2026; JevBench standings from Benchmark Heaven’s published v1.6.0 results. Our measurement ran the project’s own server on a laptop against the same items and question used for the Jev-as-a-Judge page, with the script in the repository.
When. Published October 9, 2026. Kev is versioned by Hub tag and Jev by model name; the page will say so when either moves on.
Sources
Jev is developed by TypeSafe. Kev belongs to Jared Palmer, Qwen to Alibaba, and neither is affiliated with Jev AI.