Jev AI

Model guide · open decision model · published October 9, 2026

pplx-decider

Perplexity’s open-weights 27B decision model: what its own panel shows, what the clinical claim rests on, and what we measured.

pplx-decider is the first decision model from a large AI company with the weights released: a Qwen3.8-27B fine-tune that answers typed questions with probabilities, reads images and takes 262k tokens. Its launch chart shows it edging Jev overall while losing the reasoning benchmarks, and a single-tester clinical run has been passed around as “Jev beaten”. This page draws the vendor panel so the shape is visible, quotes the claims with their sources and limits, and adds a three-way test of our own on the same 300 verdicts we used for Jev and Kev.

85.71% vs 84.51%v1 against Jev on Perplexity’s own 7,210-row panel; wins 5 benchmarks, loses 6
643 vs 628of 669 clinical cases, v1.1 against Jev, one tester, no published setup
91.0% vs 91.3%v1.1 against Jev on our 300 HaluEval verdicts; Kev-4B 81.7%
0.066 vs 0.056expected calibration error in our test, v1.1 against Jev; lower is better

What pplx-decider is

From Perplexity’s model cards and the Decisions API documentation.

Developer
Perplexity
Versions
pplx-decider-v1-27b (1 October 2026) · pplx-decider-v1.1-27b (5 October 2026)
Base model
Qwen3.8-27B, fine-tuned; about 26B parameters, 49 GiB of BF16 weights
Licence
Apache-2.0 weights on Hugging Face
Inputs
Text, JSON and images as base64 data URLs; remote image URLs are rejected
Context
262,144 input tokens; 1 to 128 questions per request
Question types
Noul · choice · score, with probabilities and no generated text
Hosted
Perplexity’s Decisions API; also served through OpenRouter
API shape
POST api.perplexity.ai/v1/decisions, Perplexity’s own format; OpenRouter’s Decisions API accepts the TypeSafe-style body
Self-hosting
One CUDA GPU with room for about 49 GiB of weights plus working memory; Python 3.12+; a 4-bit MLX conversion exists from the community

Perplexity’s panel, drawn

Eleven public benchmarks, 7,210 rows, run by Perplexity on its v1 model, on Jev over the API, and on the Qwen base it started from. Toggle a series to see where the shapes overlap.

Perplexity’s eleven-benchmark panelACCURACY · 0–100%
6080100WinoGrandeFinPhraseRAGTruthJudgeBenchBBHJevBench hardTabFactContractNLICircaBelebeleTruthfulQA
Perplexity’s own runs of all three systems on 7,210 rows; overall 85.71% · 84.51% · 74.76%.
The axes start at 0 but the rings only mark 60 to 100, where every score sits.

How to read it

The polygons are close in size, which is the overall result: 85.71% against 84.51%. The shape is the story. pplx-decider pushes out on FinancialPhraseBank, RAGTruth, TabFact, ContractNLI, Circa: evidence checking, financial sentiment, table facts, contract entailment. Jev pushes out on WinoGrande, JudgeBench, BBH, JevBench public hard, Belebele, TruthfulQA binary: common-sense and multi-step reasoning, and the hard tier of its own benchmark. The dashed Qwen base shows how much the fine-tune added on RAGTruth and FinancialPhraseBank, and how little on Circa and ContractNLI, where the base was already there.

Three caveats travel with the chart. It is the vendor’s measurement, including the Jev column. The JevBench row is the 101-item public hard subset, not the sealed benchmark on which JevBench v1.6.1 ranks Jev #5. And the Frontier, which reproduced the table, found no independent test of it.

BenchmarkRowsJevQwen3.8-27Bpplx-decider v1
WinoGrande1,00090.70%73.10%83.30%
FinancialPhraseBank99976.98%75.68%84.18%
RAGTruth1,50077.27%61.53%88.80%
JudgeBench35078.57%68.86%78.29%
BBH75094.27%72.80%82.80%
JevBench public hard10173.27%72.28%70.30%
TabFact50089.80%78.60%90.60%
ContractNLI50077.45%80.78%80.78%
Circa50084.60%87.00%89.20%
Belebele50095.00%93.20%94.00%
TruthfulQA binary51092.00%82.80%85.40%
Overall7,21084.51%74.76%85.71%

Source: perplexity-ai/pplx-decider-v1-27b model card. v1.1’s card reports Decision Index categories instead of this panel.

Decision Index 0.3, v1.1

Hugging Face’s community suite: 20% public benchmarks, 50% private tests of the same skills, 30% private tasks from new domains.

  • Tools78.88
  • Language69.45
  • Retrieval61.26
  • Knowledge48.18
  • Arts44.66
  • Overall61.56 v1: 56.4

Perplexity says v1.1 tops this board; the gain over v1 came from lifting the causal attention mask and training on more data, including tasksource. Tools and Language are its strong categories, Knowledge and Arts its weak ones, which matches the panel: strong where the answer is in the text, weaker where it has to be known or reasoned. Jev’s standing on the same index is not published on the model card, so this chart has no Jev bar.

Source: perplexity-ai/pplx-decider-v1.1-27b model card.

The clinical claim

What is actually known about “643 to 628 of 669”.

The figure comes from a post by Maziyar Panahi, a Hugging Face engineer, reporting a run of 669 medical decision cases on pplx-decider-v1.1 and Jev: 643 correct against 628, at similar speed. That is 96.1% against 93.9%, a gap of 15 cases, from one tester and one run; the dataset, the question wording and the scoring have not been published as of October 9, 2026, so the result cannot be reproduced or checked for the kind of label problems that moved other Jev scores.

Read it as a credible signal, not a verdict: it agrees in direction with the vendor panel on text-grounded tasks. “An open model completely defeated Jev in clinical decisions” is more than the numbers carry, and “System One has been open-sourced” is wrong: Perplexity released its own model, and Jev remains closed. We will link the dataset here if it is published.

Our own test: pplx-decider v1.1, Jev and Kev-4B on the same 300 verdicts

The HaluEval reference-checking task from our Jev-as-a-Judge page, replayed on v1.1 through OpenRouter. Same items, same question, same threshold.

Dataset and question300 verdicts: 150 HaluEval QA items, each with a correct and a hallucinated answer, one Noul question, “Is the candidate answer correct for the question, according only to the knowledge passage?”; identical to our Jev-as-a-Judge and Kev tests
pplx-decider sidepplx-decider-v1.1-27b through OpenRouter’s Decisions API (Perplexity-hosted), model name perplexity/pplx-decider-v1.1-27b, 2026-10-09
Jev sidejev-latest (jev-1.13.0) through TypeSafe’s API, 2026-10-07
Kev sideKev-4B, local MLX bf16 on an Apple-silicon Mac with 48 GB, 2026-10-08

Accuracyat a 0.5 threshold

pplx-decider v1.191.0%
Jev 1.13.091.3%
Kev-4B81.7%

Correct answers acceptedrecall on supported answers

pplx-decider v1.196.7%
Jev 1.13.095.3%
Kev-4B96.0%

Hallucinated answers rejectedrecall on hallucinations

pplx-decider v1.185.3%
Jev 1.13.087.3%
Kev-4B67.3%
 pplx-decider v1.1Jev 1.13.0Kev-4B
Expected calibration error0.0660.0560.125
Median latency per verdict737 ms via OpenRouter from Asia535 ms via TypeSafe from Asia559 ms on a laptop
p95 latency925 ms1088 ms636 ms

What we found

  • pplx-decider v1.1 scored 91.0%, 0.3 points below Jev’s 91.3% on this task, accepting 96.7% of correct answers and rejecting 85.3% of hallucinated ones; Kev-4B, the 4B model on a laptop, scored 81.7%.
  • Calibration: expected calibration error 0.066 for pplx-decider, 0.056 for Jev, 0.125 for Kev-4B. Verdicts pplx-decider placed in the 0.9–1.0 band were supported 93% of the time, those in the 0.0–0.1 band 2%.
  • Latency is comparable in kind: both hosted calls from a laptop in Asia, pplx-decider through OpenRouter at a 737 ms median, Jev through TypeSafe at 535 ms.

Limits of this test

  • One dataset of short question-answer pairs with model-generated hallucinations; all three scores are upper bounds for subtler production errors.
  • One question wording, written for Jev and reused verbatim; a task-shaped prompt would favour none of them in particular, but none was tuned.
  • One run each; pplx-decider resolved to the dated snapshot pplx-decider-v1.1-27b-20261006 on OpenRouter.
  • Labels are the dataset’s own; nothing was relabelled or filtered after the run. Script and all three result files are in the repository.

Call pplx-decider

Hosted by Perplexity, relayed by OpenRouter, or run from the Apache-2.0 weights.

# Through OpenRouter's Decisions API, the body our test sent (TypeSafe-style):
curl -s https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "perplexity/pplx-decider-v1.1-27b",
    "state": {"knowledge": "…", "question": "…", "candidate_answer": "…"},
    "questions": {
      "supported": {"type": "noul",
        "instructions": "Is the candidate answer correct for the question, according only to the knowledge passage?",
        "criteria": {"true": "The answer states what the passage supports.",
                     "false": "The answer contradicts or is not supported by the passage."}}
    }
  }'
# Perplexity's own endpoint: POST https://api.perplexity.ai/v1/decisions with model pplx-decider-v1.1-27b.

Perplexity’s own endpoint accepts 1 to 128 questions per request, 262,144 input tokens and images as base64 data URLs, with a 10 requests-per-second limit per organisation; the documentation notes answers can differ in the second decimal place between runs. Self-hosting needs a GPU with about 49 GiB free for the BF16 weights; the model card ships the inference code. The where to run Jev page lists the equivalent routes for Jev.

pplx-decider or Jev?

Pick pplx-decider if

  • Your decisions are grounded in supplied text: evidence checks, document and table facts, contract or policy entailment, where its panel leads.
  • You need images in the state, a 250k-token context, or weights you can run inside your own network.
  • A 27B model on an 80 GB GPU is acceptable for self-hosting.

Pick Jev if

  • The task is common-sense or multi-step reasoning, where Jev leads by 7 to 11 points on the vendor’s own panel.
  • You want an independently ranked model: Jev is #5 on JevBench v1.6.1 with calibration 90.6; pplx-decider is unmeasured there.
  • You already integrate TypeSafe’s request shape and want a pinnable hosted version with nothing to serve.

Use both when

  • You route by task: pplx-decider for grounded checks, Jev for reasoning, on the same typed questions.
  • You send the same request to both and escalate disagreements.

Jev-as-a-Judge · Kev · Clef · JevBench

pplx-decider FAQ

What is pplx-decider?

Perplexity’s decision model: a fine-tune of Qwen3.8-27B that reads a state and typed questions and returns a probability for every option, with no generated text. Version 1 shipped on 1 October 2026 and v1.1 on 5 October, both with Apache-2.0 weights on Hugging Face, a hosted Decisions API, and a listing on OpenRouter.

Did pplx-decider beat Jev?

On Perplexity’s own eleven-benchmark panel, v1 scored 85.71% to Jev’s 84.51%, winning 5 benchmarks to Jev’s 6; Jev kept the reasoning-heavy ones such as BBH and WinoGrande. A clinical test by one Hugging Face engineer scored v1.1 at 643 of 669 against Jev’s 628. In our own 300-verdict HaluEval test v1.1 scored 91% against Jev’s 91.3%. “Beat” is a fair word for some tasks and not for others; no third party has yet measured both under one method.

Is pplx-decider on JevBench?

No. JevBench v1.6.1 does not list it; Jev is #5 there at 71.5. Hugging Face’s Decision Index 0.3, a different community suite, scores v1.1 at 61.56 overall, up from 56.4 for v1, which Perplexity describes as the top of that board.

Is the Decisions API compatible with Jev requests?

Perplexity’s own endpoint, POST api.perplexity.ai/v1/decisions, uses Perplexity’s format and a 262,144-token input limit with 1 to 128 questions per request. OpenRouter serves the same model through its Decisions API, which accepts the TypeSafe-style body our test used: state, typed questions, and the model name perplexity/pplx-decider-v1.1-27b.

Can I run pplx-decider in my own environment?

Yes. The weights are Apache-2.0 and the model card ships inference code; you need a CUDA GPU with room for about 49 GiB of BF16 weights plus working memory. A community 4-bit MLX conversion exists for Apple silicon. Kev-4B is the open alternative that fits a 32 GB Mac; Clef Flash the one that fits a 39 ms budget.

Does pplx-decider read images?

Yes, as base64 data URLs in the state; remote image URLs return 400, and very large images (over 2,048 tiles) time out. Jev is text only; Clef and Jev-Omni are the other image-capable models documented on this site.

Is pplx-decider affiliated with TypeSafe?

No. It is Perplexity’s own model and API. Reports that “System One has been open-sourced” refer to Perplexity releasing its weights, not to TypeSafe opening Jev, which remains a hosted, closed model.

About this page

How it was made, so you can judge it.

Who. Jev AI operates an independent playground and API for Jev and sells access to it; pplx-decider is a competitor to the product we sell. That is why the vendor panel is reproduced in full including the rows Jev loses, the clinical claim is quoted with its author and its gaps, and our own test is published with the script. We are not affiliated with Perplexity, TypeSafe or Hugging Face.

How. Facts and the panel are from Perplexity’s model cards and API documentation as read on October 9, 2026; the clinical figures from the tester’s public post as reported; our measurement through OpenRouter’s Decisions API, on the same items and question as our Jev and Kev tests.

When. Published October 9, 2026. pplx-decider has shipped two versions in a week; the page names the snapshot we tested and will say so when a newer one lands.

Sources

Jev is developed by TypeSafe. pplx-decider belongs to Perplexity, Qwen to Alibaba, Kev to Jared Palmer; none is affiliated with Jev AI.