Model guide · open decision model · published October 9, 2026
pplx-decider
Perplexity’s open-weights 27B decision model: what its own panel shows, what the clinical claim rests on, and what we measured.
pplx-decider is the first decision model from a large AI company with the weights released: a Qwen3.8-27B fine-tune that answers typed questions with probabilities, reads images and takes 262k tokens. Its launch chart shows it edging Jev overall while losing the reasoning benchmarks, and a single-tester clinical run has been passed around as “Jev beaten”. This page draws the vendor panel so the shape is visible, quotes the claims with their sources and limits, and adds a three-way test of our own on the same 300 verdicts we used for Jev and Kev.
What pplx-decider is
From Perplexity’s model cards and the Decisions API documentation.
- Developer
- Perplexity
- Versions
- pplx-decider-v1-27b (1 October 2026) · pplx-decider-v1.1-27b (5 October 2026)
- Base model
- Qwen3.8-27B, fine-tuned; about 26B parameters, 49 GiB of BF16 weights
- Licence
- Apache-2.0 weights on Hugging Face
- Inputs
- Text, JSON and images as base64 data URLs; remote image URLs are rejected
- Context
- 262,144 input tokens; 1 to 128 questions per request
- Question types
- Noul · choice · score, with probabilities and no generated text
- Hosted
- Perplexity’s Decisions API; also served through OpenRouter
- API shape
- POST api.perplexity.ai/v1/decisions, Perplexity’s own format; OpenRouter’s Decisions API accepts the TypeSafe-style body
- Self-hosting
- One CUDA GPU with room for about 49 GiB of weights plus working memory; Python 3.12+; a 4-bit MLX conversion exists from the community
Perplexity’s panel, drawn
Eleven public benchmarks, 7,210 rows, run by Perplexity on its v1 model, on Jev over the API, and on the Qwen base it started from. Toggle a series to see where the shapes overlap.
The axes start at 0 but the rings only mark 60 to 100, where every score sits.
How to read it
The polygons are close in size, which is the overall result: 85.71% against 84.51%. The shape is the story. pplx-decider pushes out on FinancialPhraseBank, RAGTruth, TabFact, ContractNLI, Circa: evidence checking, financial sentiment, table facts, contract entailment. Jev pushes out on WinoGrande, JudgeBench, BBH, JevBench public hard, Belebele, TruthfulQA binary: common-sense and multi-step reasoning, and the hard tier of its own benchmark. The dashed Qwen base shows how much the fine-tune added on RAGTruth and FinancialPhraseBank, and how little on Circa and ContractNLI, where the base was already there.
Three caveats travel with the chart. It is the vendor’s measurement, including the Jev column. The JevBench row is the 101-item public hard subset, not the sealed benchmark on which JevBench v1.6.1 ranks Jev #5. And the Frontier, which reproduced the table, found no independent test of it.
| Benchmark | Rows | Jev | Qwen3.8-27B | pplx-decider v1 |
|---|---|---|---|---|
| WinoGrande | 1,000 | 90.70% | 73.10% | 83.30% |
| FinancialPhraseBank | 999 | 76.98% | 75.68% | 84.18% |
| RAGTruth | 1,500 | 77.27% | 61.53% | 88.80% |
| JudgeBench | 350 | 78.57% | 68.86% | 78.29% |
| BBH | 750 | 94.27% | 72.80% | 82.80% |
| JevBench public hard | 101 | 73.27% | 72.28% | 70.30% |
| TabFact | 500 | 89.80% | 78.60% | 90.60% |
| ContractNLI | 500 | 77.45% | 80.78% | 80.78% |
| Circa | 500 | 84.60% | 87.00% | 89.20% |
| Belebele | 500 | 95.00% | 93.20% | 94.00% |
| TruthfulQA binary | 510 | 92.00% | 82.80% | 85.40% |
| Overall | 7,210 | 84.51% | 74.76% | 85.71% |
Source: perplexity-ai/pplx-decider-v1-27b model card. v1.1’s card reports Decision Index categories instead of this panel.
Decision Index 0.3, v1.1
Hugging Face’s community suite: 20% public benchmarks, 50% private tests of the same skills, 30% private tasks from new domains.
Perplexity says v1.1 tops this board; the gain over v1 came from lifting the causal attention mask and training on more data, including tasksource. Tools and Language are its strong categories, Knowledge and Arts its weak ones, which matches the panel: strong where the answer is in the text, weaker where it has to be known or reasoned. Jev’s standing on the same index is not published on the model card, so this chart has no Jev bar.
The clinical claim
What is actually known about “643 to 628 of 669”.
The figure comes from a post by Maziyar Panahi, a Hugging Face engineer, reporting a run of 669 medical decision cases on pplx-decider-v1.1 and Jev: 643 correct against 628, at similar speed. That is 96.1% against 93.9%, a gap of 15 cases, from one tester and one run; the dataset, the question wording and the scoring have not been published as of October 9, 2026, so the result cannot be reproduced or checked for the kind of label problems that moved other Jev scores.
Read it as a credible signal, not a verdict: it agrees in direction with the vendor panel on text-grounded tasks. “An open model completely defeated Jev in clinical decisions” is more than the numbers carry, and “System One has been open-sourced” is wrong: Perplexity released its own model, and Jev remains closed. We will link the dataset here if it is published.
Our own test: pplx-decider v1.1, Jev and Kev-4B on the same 300 verdicts
The HaluEval reference-checking task from our Jev-as-a-Judge page, replayed on v1.1 through OpenRouter. Same items, same question, same threshold.
| Dataset and question | 300 verdicts: 150 HaluEval QA items, each with a correct and a hallucinated answer, one Noul question, “Is the candidate answer correct for the question, according only to the knowledge passage?”; identical to our Jev-as-a-Judge and Kev tests |
|---|---|
| pplx-decider side | pplx-decider-v1.1-27b through OpenRouter’s Decisions API (Perplexity-hosted), model name perplexity/pplx-decider-v1.1-27b, 2026-10-09 |
| Jev side | jev-latest (jev-1.13.0) through TypeSafe’s API, 2026-10-07 |
| Kev side | Kev-4B, local MLX bf16 on an Apple-silicon Mac with 48 GB, 2026-10-08 |
Accuracyat a 0.5 threshold
Correct answers acceptedrecall on supported answers
Hallucinated answers rejectedrecall on hallucinations
| pplx-decider v1.1 | Jev 1.13.0 | Kev-4B | |
|---|---|---|---|
| Expected calibration error | 0.066 | 0.056 | 0.125 |
| Median latency per verdict | 737 ms via OpenRouter from Asia | 535 ms via TypeSafe from Asia | 559 ms on a laptop |
| p95 latency | 925 ms | 1088 ms | 636 ms |
What we found
- pplx-decider v1.1 scored 91.0%, 0.3 points below Jev’s 91.3% on this task, accepting 96.7% of correct answers and rejecting 85.3% of hallucinated ones; Kev-4B, the 4B model on a laptop, scored 81.7%.
- Calibration: expected calibration error 0.066 for pplx-decider, 0.056 for Jev, 0.125 for Kev-4B. Verdicts pplx-decider placed in the 0.9–1.0 band were supported 93% of the time, those in the 0.0–0.1 band 2%.
- Latency is comparable in kind: both hosted calls from a laptop in Asia, pplx-decider through OpenRouter at a 737 ms median, Jev through TypeSafe at 535 ms.
Limits of this test
- One dataset of short question-answer pairs with model-generated hallucinations; all three scores are upper bounds for subtler production errors.
- One question wording, written for Jev and reused verbatim; a task-shaped prompt would favour none of them in particular, but none was tuned.
- One run each; pplx-decider resolved to the dated snapshot pplx-decider-v1.1-27b-20261006 on OpenRouter.
- Labels are the dataset’s own; nothing was relabelled or filtered after the run. Script and all three result files are in the repository.
Call pplx-decider
Hosted by Perplexity, relayed by OpenRouter, or run from the Apache-2.0 weights.
# Through OpenRouter's Decisions API, the body our test sent (TypeSafe-style):
curl -s https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "perplexity/pplx-decider-v1.1-27b",
"state": {"knowledge": "…", "question": "…", "candidate_answer": "…"},
"questions": {
"supported": {"type": "noul",
"instructions": "Is the candidate answer correct for the question, according only to the knowledge passage?",
"criteria": {"true": "The answer states what the passage supports.",
"false": "The answer contradicts or is not supported by the passage."}}
}
}'
# Perplexity's own endpoint: POST https://api.perplexity.ai/v1/decisions with model pplx-decider-v1.1-27b.Perplexity’s own endpoint accepts 1 to 128 questions per request, 262,144 input tokens and images as base64 data URLs, with a 10 requests-per-second limit per organisation; the documentation notes answers can differ in the second decimal place between runs. Self-hosting needs a GPU with about 49 GiB free for the BF16 weights; the model card ships the inference code. The where to run Jev page lists the equivalent routes for Jev.
pplx-decider or Jev?
Pick pplx-decider if
- Your decisions are grounded in supplied text: evidence checks, document and table facts, contract or policy entailment, where its panel leads.
- You need images in the state, a 250k-token context, or weights you can run inside your own network.
- A 27B model on an 80 GB GPU is acceptable for self-hosting.
Pick Jev if
- The task is common-sense or multi-step reasoning, where Jev leads by 7 to 11 points on the vendor’s own panel.
- You want an independently ranked model: Jev is #5 on JevBench v1.6.1 with calibration 90.6; pplx-decider is unmeasured there.
- You already integrate TypeSafe’s request shape and want a pinnable hosted version with nothing to serve.
Use both when
- You route by task: pplx-decider for grounded checks, Jev for reasoning, on the same typed questions.
- You send the same request to both and escalate disagreements.
Jev-as-a-Judge · Kev · Clef · JevBench
pplx-decider FAQ
What is pplx-decider?
Perplexity’s decision model: a fine-tune of Qwen3.8-27B that reads a state and typed questions and returns a probability for every option, with no generated text. Version 1 shipped on 1 October 2026 and v1.1 on 5 October, both with Apache-2.0 weights on Hugging Face, a hosted Decisions API, and a listing on OpenRouter.
Did pplx-decider beat Jev?
On Perplexity’s own eleven-benchmark panel, v1 scored 85.71% to Jev’s 84.51%, winning 5 benchmarks to Jev’s 6; Jev kept the reasoning-heavy ones such as BBH and WinoGrande. A clinical test by one Hugging Face engineer scored v1.1 at 643 of 669 against Jev’s 628. In our own 300-verdict HaluEval test v1.1 scored 91% against Jev’s 91.3%. “Beat” is a fair word for some tasks and not for others; no third party has yet measured both under one method.
Is pplx-decider on JevBench?
No. JevBench v1.6.1 does not list it; Jev is #5 there at 71.5. Hugging Face’s Decision Index 0.3, a different community suite, scores v1.1 at 61.56 overall, up from 56.4 for v1, which Perplexity describes as the top of that board.
Is the Decisions API compatible with Jev requests?
Perplexity’s own endpoint, POST api.perplexity.ai/v1/decisions, uses Perplexity’s format and a 262,144-token input limit with 1 to 128 questions per request. OpenRouter serves the same model through its Decisions API, which accepts the TypeSafe-style body our test used: state, typed questions, and the model name perplexity/pplx-decider-v1.1-27b.
Can I run pplx-decider in my own environment?
Yes. The weights are Apache-2.0 and the model card ships inference code; you need a CUDA GPU with room for about 49 GiB of BF16 weights plus working memory. A community 4-bit MLX conversion exists for Apple silicon. Kev-4B is the open alternative that fits a 32 GB Mac; Clef Flash the one that fits a 39 ms budget.
Does pplx-decider read images?
Yes, as base64 data URLs in the state; remote image URLs return 400, and very large images (over 2,048 tiles) time out. Jev is text only; Clef and Jev-Omni are the other image-capable models documented on this site.
Is pplx-decider affiliated with TypeSafe?
No. It is Perplexity’s own model and API. Reports that “System One has been open-sourced” refer to Perplexity releasing its weights, not to TypeSafe opening Jev, which remains a hosted, closed model.
About this page
How it was made, so you can judge it.
Who. Jev AI operates an independent playground and API for Jev and sells access to it; pplx-decider is a competitor to the product we sell. That is why the vendor panel is reproduced in full including the rows Jev loses, the clinical claim is quoted with its author and its gaps, and our own test is published with the script. We are not affiliated with Perplexity, TypeSafe or Hugging Face.
How. Facts and the panel are from Perplexity’s model cards and API documentation as read on October 9, 2026; the clinical figures from the tester’s public post as reported; our measurement through OpenRouter’s Decisions API, on the same items and question as our Jev and Kev tests.
When. Published October 9, 2026. pplx-decider has shipped two versions in a week; the page names the snapshot we tested and will say so when a newer one lands.
Sources
- perplexity-ai/pplx-decider-v1-27b — model card with the eleven-benchmark panel
- perplexity-ai/pplx-decider-v1.1-27b — model card with Decision Index 0.3 categories
- OpenRouter — pplx-decider v1.1 listing
- Perplexity Developers — v1.1 announcement
- Maziyar Panahi — the clinical comparison post
- The Frontier — Decisions API and panel write-up
- Hugging Face — Decision Index
- RUCAIBox/HaluEval — the QA split used in our test
Jev is developed by TypeSafe. pplx-decider belongs to Perplexity, Qwen to Alibaba, Kev to Jared Palmer; none is affiliated with Jev AI.