Jev agent · decision layer for agent loops · published October 9, 2026
Jev Agent: How Jev Works Inside an AI Agent Loop
“Jev agent” is not a product. It is an agent whose loop asks Jev the questions it used to ask an LLM: which tool, is this safe, is it done, how sure.
Jev does not plan, write or call tools. It reads the agent’s state and typed questions and returns a probability per option in about half a second. That makes it the piece of an agent loop that routes, gates, verifies and escalates, while the LLM keeps the writing and your code keeps control. This page shows where it sits, how to call it from each stack, who is already doing it, and what we measured when we gave it 440 real tool-routing decisions.
Is Jev an agent?
No. Three parts do three jobs.
| The LLM | Reads the task, plans, writes code or text, fills tool arguments. Open-ended output, seconds per call. |
|---|---|
| Jev | Answers typed questions about the current state: which tool, is this safe, is it done, how confident. A probability per option, about half a second. |
| Your code | Owns the loop: runs tools, enforces permissions, applies the thresholds, decides when to stop or ask a person. |
An agent is an LLM in a loop with tools. Most of the loop’s failures are not bad writing; they are bad decisions: the wrong tool, a destructive command waved through, a task reported finished that was not. Those are yes/no and which-one questions, and a text generator answers them slowly, expensively and without a number to branch on. Jev answers only that kind of question, which is why it fits the loop. The model guide covers the three question types; Jev vs LLM covers why generation is the wrong tool for a decision.
Four places Jev sits in the loop
Each one is a ready judge in the playground: open it, edit the criteria, export the request.
- 01
Route the next step Choice
Give Jev the task, the recent tool results and the list of tools. It returns one choice with a probability for every option, including no tool. Our measurement below is exactly this decision.
MCP tool router judge - 02
Gate a proposed action Noul
Before a shell command, a file delete or a payment runs, ask yes/no questions: does it touch production, does it delete data, is it reversible. The agent acts above a threshold and asks a person below it.
Rules checker judge - 03
Verify a completion claim Noul
When the LLM says done, Jev checks the claim against the tool output and the task. A false “done” is the most expensive agent failure; this catches it before the report.
Agent evaluation judge - 04
Guard what the agent reads Noul
Web pages, files and tool output can carry instructions aimed at the agent. One question over the text flags overrides and exfiltration before the LLM sees it as context.
Prompt injection judge
One request, four answers
A gate check before a shell command. Three Noul questions and one Choice travel in the same call; the loop branches on the numbers that come back.
Create a key on the API page. The key stays in the server process; the agent never sees it.
curl -X POST https://jev-ai.pro/api/v1/systemone \
-H "Authorization: Bearer $JEV_AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": {
"task": "Rotate the staging database password and update the app config",
"proposed_command": "psql -h db.prod.internal -c \"ALTER USER app PASSWORD ...\"",
"recent_tool_output": "Found config/database.yml with host db.prod.internal"
},
"questions": {
"touches_prod": { "type": "noul", "instructions": "Does the proposed command act on a production system?" },
"reversible": { "type": "noul", "instructions": "Can the effect of the command be undone without data loss?" },
"matches_task": { "type": "noul", "instructions": "Is the command a faithful step toward the stated task?" },
"next": { "type": "choice", "instructions": "What should the agent do now?",
"criteria": { "run": "Run the command", "ask": "Ask the user before running it", "revise": "Rewrite the command first" } }
}
}'The response carries touches_prod.noul, reversible.noul, matches_task.noul and next.choice with a probability per option. We ran this exact request on 2026-10-09: touches_prod 0.92, matches_task 0.12 because the task says staging and the command targets a production host, reversible 0.75, and for next: run 0.00, revise 0.52, ask 0.48. The loop never runs the command; it goes back to the LLM to rewrite it or to the user, which is the branch you want.
How to use Jev in your agent
Six ways in, from a curl call to a framework middleware.
Raw HTTP
POST to the systemone endpoint with your key in the Authorization header; one request can carry several questions. Works from any language, including shell hooks.
See the exampleTypeSafe’s official SDK
@typesafe-ai/sdk with baseURL pointed at Jev AI. The choice, noul and score helpers build the questions; the response is typed.
See the exampleVercel AI SDK
Vercel’s “Jev agent control” template calls Jev through AI Gateway with the SDK’s experimental_evaluate, for choosing the next evidence source and assessing a proposed action. Also works with eve, Vercel’s agent framework.
OpenLangChain / LangGraph
The langchain-typesafe package exposes Jev as a model; LangChain’s own post builds two middlewares with it, ModelRouterMiddleware (pick the model per request) and AutoModeMiddleware (gate risky tool calls before they run).
OpenClaude Code, Codex, OpenCode
Paste the setup prompt from the coding-agents page; the agent wires the API in and can call Jev from hooks, skills or its own tool loop. jevgrep adds Jev-ranked repository search.
OpenPlayground first
Build the question in the playground against a real state, read the probabilities, then export the request as curl or JSON and drop it into the loop.
OpenA loop step in code
The LLM proposes, Jev decides, your code acts. The thresholds are the whole policy, and they are numbers you can tune from logs.
// One iteration of an agent loop with Jev at the gate (Node 22+, server-side)
const JEV = 'https://jev-ai.pro/api/v1/systemone'
const headers = { authorization: `Bearer ${process.env.JEV_AI_API_KEY}`, 'content-type': 'application/json' }
async function decide(state, questions) {
const res = await fetch(JEV, { method: 'POST', headers, body: JSON.stringify({ model: 'jev-latest', state, questions }) })
if (!res.ok) throw new Error(`Jev ${res.status}`)
return (await res.json()).answers
}
async function step(task, history, tools) {
const plan = await llm.propose({ task, history, tools }) // the LLM writes the next action
const a = await decide(
{ task, proposed: plan, recent: history.slice(-3), tools: tools.map(t => ({ name: t.name, description: t.description })) },
{
tool: { type: 'choice', instructions: 'Which tool should the agent call to satisfy the task?', criteria: { ...Object.fromEntries(tools.map(t => [t.name, t.description])), none: 'No listed tool fits' } },
risky: { type: 'noul', instructions: 'Would running the proposed action delete data, spend money or touch production?' },
done: { type: 'noul', instructions: 'Is the task already complete according to the history?' },
},
)
if (a.done.noul >= 0.9) return { status: 'done' }
if (a.risky.noul >= 0.5) return { status: 'ask_user', plan } // a person confirms anything risky
if (a.tool.probabilities[a.tool.choice] < 0.7) return { status: 'ask_user', plan }
if (a.tool.choice === 'none') return { status: 'answer', plan }
return { status: 'run', tool: a.tool.choice, args: plan.args } // your code runs the tool
}Three rules do the work: stop when done is above 0.9, hand anything risky to a person, and only run a tool the model is at least 0.7 sure about. On our test, acting only above 0.9 lifted accuracy from 93.0% to 95.6% and escalated 34 of 440 decisions.
With TypeSafe’s official SDK
The same call through @typesafe-ai/sdk, pointed at Jev AI.
// Official TypeSafe SDK, published by TypeSafe — @typesafe-ai/sdk 0.6.0
// Server-side JavaScript / TypeScript using the Jev AI compatible endpoint
import { TypeSafeClient, choice } from '@typesafe-ai/sdk'
const apiKey = process.env.JEV_AI_API_KEY
if (!apiKey) throw new Error('Set JEV_AI_API_KEY on your server')
const client = new TypeSafeClient({
apiKey,
baseURL: 'https://jev-ai.pro/api', // Required, including for existing SDK users
retry: { maxRetries: 0 }, // Avoid replaying a request with an uncertain outcome
})
// The SDK appends /v1/systemone to baseURL.
const response = await client.systemOne({
model: 'jev-latest',
state: { document: 'I was charged twice. Please fix this ASAP.' },
questions: {
category: choice('What is this ticket about?', { billing: null, technical: null, other: null }),
},
})
console.log(response.answers.category.choice)Setting baseURL is required; changing only the key sends the request to TypeSafe with a key it does not recognise. The documentation covers retries, errors and limits; where to run Jev lists OpenRouter, Vercel AI Gateway, Cloudflare, Bedrock, Vertex and Azure for loops that already live there.
Cases: agents that already call Jev
Public integrations and our own judges, each with a link to the code or the measurement.
LangChain blog, October 2026
LangChain harness
Two middlewares around a LangChain agent: Jev routes each request to a cheap or capable model, and gates tool calls by risk before execution, so the generative model only runs where generation is needed.
Read itVercel template
Vercel “Jev agent control”
An AI SDK agent where Jev picks the next evidence source and assesses each proposed action; the template keeps permissions, validation and tool execution in application code, with Jev supplying the judgment.
Read itOpen source, Search1API
Jev Search
A search agent that asks Jev which engines, time range and query rewrite to use, then has Jev rank the results. Our BM25 rerank test on the same idea lifted nDCG@10 from 0.608 to 0.708.
OpenOpen source CLI
jevgrep
Repository search for coding agents: Jev scores files and excerpts for relevance to a question. The upstream project reports a 28.6% cut in agent cost on tuned SWE-bench tasks.
OpenOur measurement
Jev-as-a-Judge
Jev as the grader inside agent evals and RL reward loops: 91.3% on 300 HaluEval verdicts with calibration error 0.056, so a threshold means what it says.
OpenReady judges
Support and sales agents
Ticket triage, lead qualification and sales-call scoring are the same pattern outside coding: the agent drafts, Jev classifies and scores, your code routes.
OpenOur test: Jev picking the next tool
440 decisions from the Berkeley Function Calling Leaderboard, with a ground-truth answer for every one.
| Dataset | BFCL v4 multiple + irrelevance: 200 requests with two to four candidate tools and one correct answer, plus 240 requests where the one listed tool does not fit; Berkeley Function Calling Leaderboard v4, Apache-2.0 |
|---|---|
| Decision | One Choice question per request, “Which tool should the agent call to satisfy the user request? Pick the one tool whose purpose matches what the user asked for, or none if no listed tool fits.”, with an option per tool and a none option: “No listed tool fits the request: the agent should answer directly or ask the user instead of calling a tool.” |
| State | The user request and the tool list with each tool’s name, description and parameter descriptions; 565 input tokens per decision on average |
| Model | jev-latest (jev-1.13.0) through TypeSafe’s API, 2026-10-09; 440 decisions, every item once, no retries on a successful call |
| Scoring | The chosen option must equal the ground-truth function, or none on the irrelevance split; nothing relabelled or filtered after the run |
| Split | Decisions | Correct | What went wrong |
|---|---|---|---|
| A listed tool fits | 200 | 99.5% | 0 wrong tool, 1 declined |
| No listed tool fits | 240 | 87.5% | 30 tools called that should not have been |
| All | 440 | 93.0% | chance 39.7%; 31 misses in total |
| Act only above | Acted | Correct when acting | Escalated |
|---|---|---|---|
| 0.5 | 440 | 93.0% | 0 |
| 0.7 | 428 | 94.4% | 12 |
| 0.8 | 421 | 94.8% | 19 |
| 0.9 | 406 | 95.6% | 34 |
What we found
- When one of the listed tools was right, Jev chose it 99.5% of the time across two-, three- and four-tool menus, and never chose a wrong tool; the one miss was a decline.
- Knowing when not to call a tool is the harder half: 87.5%. The 30 false calls are mostly near-fits, such as a projectile-range tool for a car-launch question or a heat calculator for warming water, where the dataset’s label is “no tool” but the tool would compute something.
- The probability is honest. Decisions in the 0.9–1.0 band were right 96% of the time; the 19 decisions below 0.8 were right 52.6% of the time. Mean probability on correct choices 0.985, on wrong ones 0.861.
- Latency: 472 ms median and 536 ms p95 per decision with the whole tool list in the state, 565 input tokens on average, at concurrency 4 from Asia.
Examples of misses
- “How far will a car travel in time 't' when launched with velocity 'v' at an angle 'theta'?” chose
calculate_projectile_rangeat 0.91; expectednone - “What is the magnetic field at a point located at distance 'r' from a wire carrying current 'I'?” chose
magnetic_field_intensityat 0.85; expectednone - “What will be the energy needed to increase the temperature of 3 kg of water by 4 degrees Celsius?” chose
calculate_heatat 1; expectednone - “What is the gene sequence for evolutionary changes in whales?” chose
gene_sequencerat 0.94; expectednone
Limits of this test
- BFCL’s irrelevance labels are the benchmark’s; several of the misses are defensible calls, and a production tool list would carry a no-tool policy in the description.
- One question wording, not tuned; one run; hosted Jev only. The script accepts a base URL and model, so Kev or pplx-decider can be run on the same items.
- Routing is one of the four decisions. Our judge test covers verification; the gate and guard decisions are measured only by the judges linked above.
Jev agent FAQ
Is Jev an AI agent?
No. Jev is a decision model: it reads a state and typed questions and returns a probability per option, with no generated text. An agent is a loop that an LLM drives and your code runs; Jev is the component the loop calls for routing, gating, verification and confidence. The phrase “Jev agent” usually means an agent that uses Jev that way.
Can Jev run an agent on its own?
Not usefully. It cannot plan, write code, fill tool arguments or produce text. It can pick which of several planned actions to take, say whether an action is safe, and say whether the work is done, which is the part of the loop that most often fails silently.
How accurate is Jev at picking the next tool?
In our test on 440 BFCL decisions it chose correctly 93.0% of the time: 99.5% when one of the listed tools was right, and 87.5% at declining when none fit. Acting only above probability 0.9 raised accuracy to 95.6% on 406 decisions and sent 34 to a person.
Where should the threshold be?
Read it off the calibration table for your own decisions. On this test the probability Jev put on its choice tracked how often it was right, with expected calibration error 0.046, so 0.9 means roughly nine in ten. Start at 0.9 for irreversible actions and 0.7 for routing, then move it with data.
Does Jev replace the LLM in my agent?
No. Keep Claude, GPT, Gemini or an open model for planning and writing; put Jev at the decision points. Each Jev call is a few hundred milliseconds and returns a number the loop can branch on, which is what the LLM call was being asked to do at those points.
Which frameworks support Jev?
Anything that can make an HTTP call. Named integrations: TypeSafe’s SDKs, Vercel AI SDK through AI Gateway, LangChain’s langchain-typesafe package, OpenRouter, and the coding agents on our setup page. Jev AI exposes the same request shape at its own endpoint.
Is there a Jev MCP server?
Jev AI does not publish one, and most loops do not need one: the agent’s host code calls Jev directly between steps. The MCP tool router judge shows the reverse, Jev choosing among MCP tools.
What about open models such as Kev or pplx-decider?
They take the same place in the loop and the same request shape. Kev-4B runs on a laptop and scored 81.7% on our judge test against Jev’s 91.3%; pplx-decider scored 91.0% and ships Apache-2.0 weights. Both have guides on this site.
About this page
Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and publishes measurements of it. We are not affiliated with TypeSafe, Vercel, LangChain or the BFCL authors.
How. The routing test used the public BFCL v4 multiple and irrelevance files, one Choice question per request, through TypeSafe’s API on 2026-10-09; the script and the full result file, including every miss, are in our repository. Integration descriptions are from the linked vendor pages as read on October 9, 2026.
When. Published October 9, 2026. Numbers are tied to jev-1.13.0 and will be rerun when the model changes.
Sources
TypeSafe, Vercel, LangChain and the Berkeley Function Calling Leaderboard are the work of their respective owners and are not affiliated with Jev AI.