Jev AI

Model guide · open decision model · published October 11, 2026

Jeeves

PostHog’s open 9B decision model that thinks before it answers. What it claims against Jev, and what it did on our machine.

Jeeves takes the idea behind Jev, typed questions answered with probabilities, and adds a reasoning step: before choosing, it writes a short chain of thought, sped up by a diffusion drafter. PostHog says this beats Jev on out-of-domain tasks and on JevBench’s public items. The weights, training code and data are MIT-licensed. We downloaded the FP8 build, served it on a 48 GB Mac, and ran it with and without thinking on the same 300 judging verdicts and 440 tool-routing decisions we used for Jev.

0.935 vs 0.866Jeeves against Jev on JevBench’s 231 public items, PostHog’s run with thinking
87.7% vs 91.3%our 300 HaluEval verdicts, Jeeves thinking against hosted Jev
93.0% vs 93.0%our 440 BFCL tool-routing decisions, same pair
1.2 s · 14.4 smedian per verdict on our Mac, one at a time, without and with thinking

Jeeves at a glance

DeveloperPostHog
Released29 September 2026 (weights on Hugging Face as PostHog/jeeves and PostHog/jeeves-fp8)
Base modelQwen3.5-9B with a LoRA and a pointer head that reads option probabilities
TrainingSupervised fine-tuning, then CISPO reinforcement learning so it reasons before it decides; full training code and train/dev/test data published
Speed-upA block-4 diffusion drafter for speculative decoding of the reasoning chain
LicenceMIT
APIJev-compatible POST /v1/systemone with noul, choice and score questions; a drop-in Python SDK replacement
Weights21 GB in bf16, 11.5 GB with FP8 linear layers
HardwareA CUDA GPU (FP8 needs compute capability 8.9+), or an Apple silicon Mac with 48 GB or more
Inspired byKev, the open Jev alternative from Jared Palmer

What PostHog reports

Thinking on, greedy, 2,560-token cap. The Jev and Kev columns are numbers PostHog took from Kev’s published results, not new runs.

Test overall (out-of-domain and held-out)

  • Jeeves0.889
  • Jev0.857
  • Kev0.822

Transfer overall (MMLU-Pro and buried state)

  • Jeeves0.746
  • Jev0.800
  • Kev0.579

JevBench overall (231 public items)

  • Jeeves0.935
  • Jev0.866
  • Kev0.715

JevBench hard (111 public items)

  • Jeeves0.865
  • Jev0.730
  • Kev0.451
TaskKev-9BJevJeeves
QNLI0.9250.9250.913
SciQ0.9630.9880.991
TweetEval offensive0.7750.8130.813
PAWS0.7630.7880.875
MMLU0.7380.9000.793
Emotion0.6000.5880.647
Held-out rule structures0.8960.8851.000
Contrastive policies0.9000.9631.000
MMLU-Pro (10-way)0.5150.8400.739
Buried state0.7400.7000.759

Jeeves leads on 6 of the 10 tasks, mostly rule following, paraphrase and policies, where reasoning about the stated criteria pays. Jev leads on 3, including the knowledge-heavy MMLU and MMLU-Pro, where a 9B base knows less. Without thinking the same checkpoint scores 0.804 on PostHog’s test split (2,962 items), against 0.840 with it. JevBench expected calibration error on public items: Jev 0.049, Jeeves 0.037. Unknowable questions answered at p ≥ 0.9 (lower is better): Kev-9B 0.000, Jev 0.090, Jeeves 0.055. No Kev-9B JevBench result is published; PostHog’s Kev JevBench figures are Kev-8B (Qwen3).

Our test: Jeeves on a Mac, against hosted Jev

The same two tests we ran on Jev, Kev and the Ollama models. Nothing tuned for any model.

Model and modeJudge accuracyAccepts · rejectsJudge ECERouting accuracyRight tool · declinesMedian per verdict
Jeeves, thinking (max_think 768, nothink_threshold 0.9)87.7%93.3% · 82.0%0.07393.0%99.0% · 87.9%14.4 s alone · 46.4 s at 4 concurrent
Jeeves, no thinking85.0%93.3% · 76.7%0.09491.6%99.0% · 85.4%1.2 s
Jev (jev-1.13.0), hosted91.3%95.3% · 87.3%0.05693.0%99.5% · 87.5%535 ms over the network
Kev-4B, local MLX81.7%96.0% · 67.3%0.125——559 ms

What we found

  • Thinking helped, but did not close the gap as a judge. On 300 HaluEval verdicts Jeeves scored 87.7% with thinking and 85.0% without, against Jev’s 91.3%. The gain came from catching wrong answers: 82.0% rejected with thinking against 76.7% without (Jev 87.3%).
  • On tool routing Jeeves with thinking tied Jev at 93.0%, and was the best calibrated model we have measured on this test (ECE 0.02). Without thinking it scored 91.6%. Thinking mostly improved declining when no tool fits: 87.9% against 85.4%.
  • nothink_threshold does what PostHog says: Jeeves reasoned on 42% of judging verdicts (about 648 tokens each) and only 23% of routing decisions (about 190 tokens), answering the confident ones directly.
  • On a Mac, thinking is slow. One request at a time, a thinking verdict took a 14.4 s median and 40 s at p95 on a 40-verdict sample; without thinking a verdict took 1.2 s. PostHog’s H100 figures are several times faster.
  • PostHog’s headline lead is on JevBench’s public items, which we did not rerun. On our two tests, which neither model was tuned for, Jeeves matches Jev on routing and trails it by a few points on judging.

How we tested

  • Judging: 150 HaluEval QA items, a correct and a hallucinated answer each, 300 Noul verdicts; the question wording is the one from our Jev-as-a-Judge test.
  • Routing: 200 BFCL requests with one right tool and 240 where no tool fits, one Choice each, as in our Jev agent test.
  • Jeeves: PostHog/jeeves-fp8 at commit 3f948de, served with --precision fp8 --max-rows 4 --max-len 4096 on an Apple M5 Pro MacBook with 48 GB. “Thinking” is PostHog’s balanced setting, max_think 768 and nothink_threshold 0.9, run 4 requests at a time so the server’s 4 rows stay busy; “no thinking” is think: false, one request at a time. Latency under load is per request, not per batch.
  • Limits: two datasets, one run each, one wording written for Jev. JevBench itself, where PostHog reports its biggest lead, was not rerun here.

What thinking looks like

One verdict from our run, without and with thinking.

Question: Which Oscar-nominated film was written by the screenwriter who wrote a 1991 romantic drama based upon a screenplay by Sooni Taraporevala?

Candidate answer: “The Namesake”, which is wrong: the passage calls a different film Oscar-nominated.

  • No thinking0.85probability the answer is correct: accepted, wrongly
  • Thinking, 768 tokens, 29.4 s0.49below 0.5: rejected, correctly
The knowledge states she is best known as the screenwriter of "Mississippi Masala", "The Namesake" and Oscar-nominated "Salaam Bombay". Therefore, "Salaam Bombay" is the Oscar-nominated film written by her. "The Namesake" is mentioned as a film she wrote, but the text explicitly labels "Salaam Bombay" as the "Oscar-nominated" one. Does the text say "The Namesake" is Oscar-nominated? No.

Excerpt of the reasoning Jeeves returned with return_reasoning, lightly trimmed. Note the final probability sits just under 0.5: the reasoning found the problem, but a threshold of 0.5 is doing a lot of work here. Jev returns no reasoning to read.

Thinking: the speed–accuracy dial

PostHog’s measurements on 325 dev questions, one H100, FP8.

SettingAccuracyMean reasoning tokensLatency
Full thinking0.8251,1383.3 s median, 17.1 s p90
max_think 768, nothink_threshold 0.90.8063442.0 s median, 5.6 s p90
No thinking0.7750about 0.3 s

The options travel in the request, so one server can answer quick gates without thinking and slow, important decisions with it. nothink_threshold is the clever one: Jeeves answers immediately when its first pass is confident and reasons only when it is not, which is why its median stays low while hard items get more time.

Run it yourself

Exactly what we ran, on Apple silicon with 48 GB.

git clone https://github.com/PostHog/jeeves && cd jeeves
uv venv --python 3.12 && uv pip install -r requirements.txt
hf download PostHog/jeeves-fp8 --local-dir jeeves-fp8

# Apple silicon, 48 GB: FP8 weights and smaller caches so it does not swap
python -m inference.serve --model jeeves-fp8 \
  --drafter jeeves-fp8/drafter_k4.safetensors \
  --precision fp8 --max-rows 4 --max-len 4096 --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
  "questions": {
    "department": { "type": "choice", "instructions": "Which team should handle this?",
                    "criteria": { "returns": "Exchanges and wrong items", "shipping": "Delays",
                                  "billing": "Charges and payments" } },
    "escalate":   { "type": "noul", "instructions": "Does this need urgent human attention?" }
  },
  "options": { "max_think": 768, "nothink_threshold": 0.9, "return_reasoning": true }
}'

No API key is needed locally. The sdk/ folder is a drop-in for TypeSafe’s Python SDK that defaults to http://127.0.0.1:8009. On a CUDA GPU with compute capability 8.9 or higher, FP8 runs with Triton kernels; PostHog reports about 0.3 s per request without thinking on an H100.

Jeeves or Jev?

Pick Jeeves if

  • Your decisions follow written rules or policies, where its reasoning step helps most.
  • You need to see why: return_reasoning gives the chain behind every answer.
  • Data must stay on your hardware, or you want to fine-tune on your own labels.
  • A few seconds per hard decision is acceptable.

Pick Jev if

  • Latency must stay under a second on every call, not just the median.
  • The task leans on world knowledge, where Jev leads MMLU and MMLU-Pro on PostHog’s own table.
  • You want a hosted, pinned model with nothing to serve. Try it in the playground.

Related

Kev, the model Jeeves builds on · Jev-style models on Ollama · pplx-decider · all local options · JevBench

Jeeves FAQ

What is Jeeves?

An open decision model from PostHog, released on 29 September 2026 under the MIT licence. It is Qwen3.5-9B with a LoRA and a pointer head, trained with supervised fine-tuning and then CISPO reinforcement learning to write a short chain of reasoning before choosing. It answers Jev-style noul, choice and score questions through a Jev-compatible /v1/systemone API.

Does Jeeves beat Jev?

On PostHog’s own tables, yes in places: 0.889 against 0.857 on its out-of-domain test split and 0.935 against 0.866 on JevBench’s 231 public items, while Jev keeps MMLU and MMLU-Pro. In our test on 300 HaluEval verdicts, Jeeves scored 87.7% with thinking and 85.0% without, against Jev’s 91.3%. Results depend heavily on the task and on whether thinking is on.

How fast is Jeeves?

On our Apple M5 Pro, one request at a time, a verdict took a 1.2 s median without thinking and 14.4 s with PostHog’s balanced setting (max_think 768, nothink_threshold 0.9), because it reasons whenever its first answer is not confident. PostHog reports 0.3 s without thinking and 2.0 s with that setting on an H100.

Can I run Jeeves on a Mac?

Yes, on Apple silicon with 48 GB or more. The FP8 weights take 11.5 GB; PostHog recommends --max-rows 4 --max-len 4096 so the caches fit without swapping. We ran it that way on a 48 GB M5 Pro.

Is Jeeves on Ollama?

Not as of 11 October 2026. It ships its own Python server that speaks the same /v1/systemone API Ollama uses, so clients written for Jev or Ollama can call it at localhost:8009.

Does Jeeves work with OpenClaw?

Not with the official TypeSafe plugin as of 11 October 2026. In our test, Jeeves with thinking on ran past OpenClaw’s 30-second limit for a three-question batch on a Mac, and with thinking off its answers were rejected because the usage block includes reasoning_tokens, a field the plugin does not accept. Kev works with the same plugin.

Can I see why Jeeves chose an answer?

Yes. With return_reasoning set, each answer comes with the reasoning text that produced it. Jev returns probabilities only. For auditing a decision after the fact, that is the clearest practical difference between them.

Can I fine-tune Jeeves?

Yes. PostHog published the full training code and the train, dev and test data, and an export path that fuses your own checkpoint and serves it with a drafter. Jev cannot be fine-tuned by users.

About this page

Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and publishes measurements of it and its alternatives. We are not affiliated with PostHog or TypeSafe.

How. Model facts and PostHog’s tables are from the Jeeves README and Hugging Face model card as read on October 11, 2026. Our numbers come from the scripts and result files in our repository, run on 2026-10-11.

When. Published October 11, 2026. We will rerun when PostHog ships a new checkpoint.

Sources

Jeeves is PostHog’s; Jev is TypeSafe AI’s. Neither is affiliated with Jev AI.