Jev AI

Jev agent · decision layer for agent loops · published October 9, 2026

Jev Agent: How Jev Works Inside an AI Agent Loop

“Jev agent” is not a product. It is an agent whose loop asks Jev the questions it used to ask an LLM: which tool, is this safe, is it done, how sure.

Jev does not plan, write or call tools. It reads the agent’s state and typed questions and returns a probability per option in about half a second. That makes it the piece of an agent loop that routes, gates, verifies and escalates, while the LLM keeps the writing and your code keeps control. This page shows where it sits, how to call it from each stack, who is already doing it, and what we measured when we gave it 440 real tool-routing decisions.

93.0%correct next-tool decisions on 440 BFCL requests; chance 39.7%
99.5% · 87.5%picks the right tool when one fits · declines when none does
472 msmedian per decision from Asia, 536 ms p95, with the tool list in the state
0.046expected calibration error: the probability on the chosen tool is a usable threshold

Is Jev an agent?

No. Three parts do three jobs.

The LLMReads the task, plans, writes code or text, fills tool arguments. Open-ended output, seconds per call.
JevAnswers typed questions about the current state: which tool, is this safe, is it done, how confident. A probability per option, about half a second.
Your codeOwns the loop: runs tools, enforces permissions, applies the thresholds, decides when to stop or ask a person.

An agent is an LLM in a loop with tools. Most of the loop’s failures are not bad writing; they are bad decisions: the wrong tool, a destructive command waved through, a task reported finished that was not. Those are yes/no and which-one questions, and a text generator answers them slowly, expensively and without a number to branch on. Jev answers only that kind of question, which is why it fits the loop. The model guide covers the three question types; Jev vs LLM covers why generation is the wrong tool for a decision.

Four places Jev sits in the loop

Each one is a ready judge in the playground: open it, edit the criteria, export the request.

  1. 01

    Route the next step Choice

    Give Jev the task, the recent tool results and the list of tools. It returns one choice with a probability for every option, including no tool. Our measurement below is exactly this decision.

    MCP tool router judge
  2. 02

    Gate a proposed action Noul

    Before a shell command, a file delete or a payment runs, ask yes/no questions: does it touch production, does it delete data, is it reversible. The agent acts above a threshold and asks a person below it.

    Rules checker judge
  3. 03

    Verify a completion claim Noul

    When the LLM says done, Jev checks the claim against the tool output and the task. A false “done” is the most expensive agent failure; this catches it before the report.

    Agent evaluation judge
  4. 04

    Guard what the agent reads Noul

    Web pages, files and tool output can carry instructions aimed at the agent. One question over the text flags overrides and exfiltration before the LLM sees it as context.

    Prompt injection judge

One request, four answers

A gate check before a shell command. Three Noul questions and one Choice travel in the same call; the loop branches on the numbers that come back.

Create a key on the API page. The key stays in the server process; the agent never sees it.

curl -X POST https://jev-ai.pro/api/v1/systemone \
  -H "Authorization: Bearer $JEV_AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": {
      "task": "Rotate the staging database password and update the app config",
      "proposed_command": "psql -h db.prod.internal -c \"ALTER USER app PASSWORD ...\"",
      "recent_tool_output": "Found config/database.yml with host db.prod.internal"
    },
    "questions": {
      "touches_prod": { "type": "noul", "instructions": "Does the proposed command act on a production system?" },
      "reversible":   { "type": "noul", "instructions": "Can the effect of the command be undone without data loss?" },
      "matches_task": { "type": "noul", "instructions": "Is the command a faithful step toward the stated task?" },
      "next": { "type": "choice", "instructions": "What should the agent do now?",
                "criteria": { "run": "Run the command", "ask": "Ask the user before running it", "revise": "Rewrite the command first" } }
    }
  }'

The response carries touches_prod.noul, reversible.noul, matches_task.noul and next.choice with a probability per option. We ran this exact request on 2026-10-09: touches_prod 0.92, matches_task 0.12 because the task says staging and the command targets a production host, reversible 0.75, and for next: run 0.00, revise 0.52, ask 0.48. The loop never runs the command; it goes back to the LLM to rewrite it or to the user, which is the branch you want.

How to use Jev in your agent

Six ways in, from a curl call to a framework middleware.

Raw HTTP

POST to the systemone endpoint with your key in the Authorization header; one request can carry several questions. Works from any language, including shell hooks.

See the example

TypeSafe’s official SDK

@typesafe-ai/sdk with baseURL pointed at Jev AI. The choice, noul and score helpers build the questions; the response is typed.

See the example

Vercel AI SDK

Vercel’s “Jev agent control” template calls Jev through AI Gateway with the SDK’s experimental_evaluate, for choosing the next evidence source and assessing a proposed action. Also works with eve, Vercel’s agent framework.

Open

LangChain / LangGraph

The langchain-typesafe package exposes Jev as a model; LangChain’s own post builds two middlewares with it, ModelRouterMiddleware (pick the model per request) and AutoModeMiddleware (gate risky tool calls before they run).

Open

Claude Code, Codex, OpenCode

Paste the setup prompt from the coding-agents page; the agent wires the API in and can call Jev from hooks, skills or its own tool loop. jevgrep adds Jev-ranked repository search.

Open

Playground first

Build the question in the playground against a real state, read the probabilities, then export the request as curl or JSON and drop it into the loop.

Open

A loop step in code

The LLM proposes, Jev decides, your code acts. The thresholds are the whole policy, and they are numbers you can tune from logs.

// One iteration of an agent loop with Jev at the gate (Node 22+, server-side)
const JEV = 'https://jev-ai.pro/api/v1/systemone'
const headers = { authorization: `Bearer ${process.env.JEV_AI_API_KEY}`, 'content-type': 'application/json' }

async function decide(state, questions) {
  const res = await fetch(JEV, { method: 'POST', headers, body: JSON.stringify({ model: 'jev-latest', state, questions }) })
  if (!res.ok) throw new Error(`Jev ${res.status}`)
  return (await res.json()).answers
}

async function step(task, history, tools) {
  const plan = await llm.propose({ task, history, tools })           // the LLM writes the next action
  const a = await decide(
    { task, proposed: plan, recent: history.slice(-3), tools: tools.map(t => ({ name: t.name, description: t.description })) },
    {
      tool: { type: 'choice', instructions: 'Which tool should the agent call to satisfy the task?', criteria: { ...Object.fromEntries(tools.map(t => [t.name, t.description])), none: 'No listed tool fits' } },
      risky: { type: 'noul', instructions: 'Would running the proposed action delete data, spend money or touch production?' },
      done: { type: 'noul', instructions: 'Is the task already complete according to the history?' },
    },
  )
  if (a.done.noul >= 0.9) return { status: 'done' }
  if (a.risky.noul >= 0.5) return { status: 'ask_user', plan }         // a person confirms anything risky
  if (a.tool.probabilities[a.tool.choice] < 0.7) return { status: 'ask_user', plan }
  if (a.tool.choice === 'none') return { status: 'answer', plan }
  return { status: 'run', tool: a.tool.choice, args: plan.args }       // your code runs the tool
}

Three rules do the work: stop when done is above 0.9, hand anything risky to a person, and only run a tool the model is at least 0.7 sure about. On our test, acting only above 0.9 lifted accuracy from 93.0% to 95.6% and escalated 34 of 440 decisions.

With TypeSafe’s official SDK

The same call through @typesafe-ai/sdk, pointed at Jev AI.

// Official TypeSafe SDK, published by TypeSafe — @typesafe-ai/sdk 0.6.0
// Server-side JavaScript / TypeScript using the Jev AI compatible endpoint
import { TypeSafeClient, choice } from '@typesafe-ai/sdk'

const apiKey = process.env.JEV_AI_API_KEY
if (!apiKey) throw new Error('Set JEV_AI_API_KEY on your server')

const client = new TypeSafeClient({
  apiKey,
  baseURL: 'https://jev-ai.pro/api', // Required, including for existing SDK users
  retry: { maxRetries: 0 }, // Avoid replaying a request with an uncertain outcome
})

// The SDK appends /v1/systemone to baseURL.
const response = await client.systemOne({
  model: 'jev-latest',
  state: { document: 'I was charged twice. Please fix this ASAP.' },
  questions: {
    category: choice('What is this ticket about?', { billing: null, technical: null, other: null }),
  },
})
console.log(response.answers.category.choice)

Setting baseURL is required; changing only the key sends the request to TypeSafe with a key it does not recognise. The documentation covers retries, errors and limits; where to run Jev lists OpenRouter, Vercel AI Gateway, Cloudflare, Bedrock, Vertex and Azure for loops that already live there.

Cases: agents that already call Jev

Public integrations and our own judges, each with a link to the code or the measurement.

LangChain blog, October 2026

LangChain harness

Two middlewares around a LangChain agent: Jev routes each request to a cheap or capable model, and gates tool calls by risk before execution, so the generative model only runs where generation is needed.

Read it

Vercel template

Vercel “Jev agent control”

An AI SDK agent where Jev picks the next evidence source and assesses each proposed action; the template keeps permissions, validation and tool execution in application code, with Jev supplying the judgment.

Read it

Open source, Search1API

Jev Search

A search agent that asks Jev which engines, time range and query rewrite to use, then has Jev rank the results. Our BM25 rerank test on the same idea lifted nDCG@10 from 0.608 to 0.708.

Open

Open source CLI

jevgrep

Repository search for coding agents: Jev scores files and excerpts for relevance to a question. The upstream project reports a 28.6% cut in agent cost on tuned SWE-bench tasks.

Open

Our measurement

Jev-as-a-Judge

Jev as the grader inside agent evals and RL reward loops: 91.3% on 300 HaluEval verdicts with calibration error 0.056, so a threshold means what it says.

Open

Ready judges

Support and sales agents

Ticket triage, lead qualification and sales-call scoring are the same pattern outside coding: the agent drafts, Jev classifies and scores, your code routes.

Open

Our test: Jev picking the next tool

440 decisions from the Berkeley Function Calling Leaderboard, with a ground-truth answer for every one.

DatasetBFCL v4 multiple + irrelevance: 200 requests with two to four candidate tools and one correct answer, plus 240 requests where the one listed tool does not fit; Berkeley Function Calling Leaderboard v4, Apache-2.0
DecisionOne Choice question per request, “Which tool should the agent call to satisfy the user request? Pick the one tool whose purpose matches what the user asked for, or none if no listed tool fits.”, with an option per tool and a none option: “No listed tool fits the request: the agent should answer directly or ask the user instead of calling a tool.”
StateThe user request and the tool list with each tool’s name, description and parameter descriptions; 565 input tokens per decision on average
Modeljev-latest (jev-1.13.0) through TypeSafe’s API, 2026-10-09; 440 decisions, every item once, no retries on a successful call
ScoringThe chosen option must equal the ground-truth function, or none on the irrelevance split; nothing relabelled or filtered after the run
SplitDecisionsCorrectWhat went wrong
A listed tool fits20099.5%0 wrong tool, 1 declined
No listed tool fits24087.5%30 tools called that should not have been
All44093.0%chance 39.7%; 31 misses in total
Act only aboveActedCorrect when actingEscalated
0.544093.0%0
0.742894.4%12
0.842194.8%19
0.940695.6%34

What we found

  • When one of the listed tools was right, Jev chose it 99.5% of the time across two-, three- and four-tool menus, and never chose a wrong tool; the one miss was a decline.
  • Knowing when not to call a tool is the harder half: 87.5%. The 30 false calls are mostly near-fits, such as a projectile-range tool for a car-launch question or a heat calculator for warming water, where the dataset’s label is “no tool” but the tool would compute something.
  • The probability is honest. Decisions in the 0.9–1.0 band were right 96% of the time; the 19 decisions below 0.8 were right 52.6% of the time. Mean probability on correct choices 0.985, on wrong ones 0.861.
  • Latency: 472 ms median and 536 ms p95 per decision with the whole tool list in the state, 565 input tokens on average, at concurrency 4 from Asia.

Examples of misses

  • “How far will a car travel in time 't' when launched with velocity 'v' at an angle 'theta'?” chose calculate_projectile_range at 0.91; expected none
  • “What is the magnetic field at a point located at distance 'r' from a wire carrying current 'I'?” chose magnetic_field_intensity at 0.85; expected none
  • “What will be the energy needed to increase the temperature of 3 kg of water by 4 degrees Celsius?” chose calculate_heat at 1; expected none
  • “What is the gene sequence for evolutionary changes in whales?” chose gene_sequencer at 0.94; expected none

Limits of this test

  • BFCL’s irrelevance labels are the benchmark’s; several of the misses are defensible calls, and a production tool list would carry a no-tool policy in the description.
  • One question wording, not tuned; one run; hosted Jev only. The script accepts a base URL and model, so Kev or pplx-decider can be run on the same items.
  • Routing is one of the four decisions. Our judge test covers verification; the gate and guard decisions are measured only by the judges linked above.

Jev agent FAQ

Is Jev an AI agent?

No. Jev is a decision model: it reads a state and typed questions and returns a probability per option, with no generated text. An agent is a loop that an LLM drives and your code runs; Jev is the component the loop calls for routing, gating, verification and confidence. The phrase “Jev agent” usually means an agent that uses Jev that way.

Can Jev run an agent on its own?

Not usefully. It cannot plan, write code, fill tool arguments or produce text. It can pick which of several planned actions to take, say whether an action is safe, and say whether the work is done, which is the part of the loop that most often fails silently.

How accurate is Jev at picking the next tool?

In our test on 440 BFCL decisions it chose correctly 93.0% of the time: 99.5% when one of the listed tools was right, and 87.5% at declining when none fit. Acting only above probability 0.9 raised accuracy to 95.6% on 406 decisions and sent 34 to a person.

Where should the threshold be?

Read it off the calibration table for your own decisions. On this test the probability Jev put on its choice tracked how often it was right, with expected calibration error 0.046, so 0.9 means roughly nine in ten. Start at 0.9 for irreversible actions and 0.7 for routing, then move it with data.

Does Jev replace the LLM in my agent?

No. Keep Claude, GPT, Gemini or an open model for planning and writing; put Jev at the decision points. Each Jev call is a few hundred milliseconds and returns a number the loop can branch on, which is what the LLM call was being asked to do at those points.

Which frameworks support Jev?

Anything that can make an HTTP call. Named integrations: TypeSafe’s SDKs, Vercel AI SDK through AI Gateway, LangChain’s langchain-typesafe package, OpenRouter, and the coding agents on our setup page. Jev AI exposes the same request shape at its own endpoint.

Is there a Jev MCP server?

Jev AI does not publish one, and most loops do not need one: the agent’s host code calls Jev directly between steps. The MCP tool router judge shows the reverse, Jev choosing among MCP tools.

What about open models such as Kev or pplx-decider?

They take the same place in the loop and the same request shape. Kev-4B runs on a laptop and scored 81.7% on our judge test against Jev’s 91.3%; pplx-decider scored 91.0% and ships Apache-2.0 weights. Both have guides on this site.

About this page

Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and publishes measurements of it. We are not affiliated with TypeSafe, Vercel, LangChain or the BFCL authors.

How. The routing test used the public BFCL v4 multiple and irrelevance files, one Choice question per request, through TypeSafe’s API on 2026-10-09; the script and the full result file, including every miss, are in our repository. Integration descriptions are from the linked vendor pages as read on October 9, 2026.

When. Published October 9, 2026. Numbers are tied to jev-1.13.0 and will be rerun when the model changes.

Sources

TypeSafe, Vercel, LangChain and the Berkeley Function Calling Leaderboard are the work of their respective owners and are not affiliated with Jev AI.