Jev AI

Local models · Ollama · published October 11, 2026

Jev on Ollama

You cannot pull Jev itself. Since Ollama 0.35 you can run open models behind Jev’s exact API, on your own machine. We measured how close they get.

Short answer: TypeSafe’s Jev is closed and is not on Ollama. On 29 September 2026 Ollama added a /v1/systemone endpoint that speaks Jev’s request format, launched with three open decision models, nimble (9B, Bespoke Labs) and tev1 in 4B and 0.8B (Together AI), and has since added Cloudflare’s clef and clef-flash and Convai’s laya. Code written for Jev runs against them with a changed base URL. We pulled 4 of them and ran them on the same 300 judging verdicts and 440 tool-routing decisions we used for hosted Jev.

86.7% vs 91.3%best local judge (nimble) against hosted Jev on 300 HaluEval verdicts
93.0% vs 93.0%best local router (tev1 (4B)) against Jev on 440 BFCL decisions
18 msmedian verdict for laya on the laptop; Jev 535 ms over the network
0.35+Ollama version with /v1/systemone; we used 0.40.2

Set it up in two minutes

Upgrade Ollama, pull a decision model, send the same JSON you send Jev.

brew upgrade ollama        # or download 0.35+ from ollama.com
ollama pull nimble          # 9B, Bespoke Labs, Apache-2.0
ollama pull tev1            # 4B, Together AI
ollama pull tev1:0.8b       # 0.8B, Together AI
curl http://localhost:11434/v1/systemone -d '{
  "model": "nimble",
  "state": { "ticket": "I was charged twice. Please refund the extra payment." },
  "questions": {
    "team":    { "type": "choice", "instructions": "Which team should handle this ticket?",
                 "criteria": { "billing": "Payments and refunds", "technical": "Bugs", "other": "None of the above" } },
    "refund":  { "type": "noul",   "instructions": "Does the customer explicitly ask for a refund?" },
    "urgency": { "type": "score",  "instructions": "How urgent is this ticket?",
                 "criteria": ["Routine", "Soon", "Urgent"] }
  }
}'

The response has the same shape as Jev’s: choice with a probability per option and a confidence, noul as the probability of yes, and score as a probability-weighted level with its legend. Ollama builds each model’s prompt from your questions; there is no prompt to write.

Switch between local and hosted with two variables

TypeSafe’s SDK reads its address and key from the environment.

# TypeSafe's Python SDK, pointed at Ollama (the key is ignored but required)
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama
export TYPESAFE_DEFAULT_MODEL=nimble

# Switch back to hosted Jev through Jev AI by changing two variables
export TYPESAFE_BASE_URL=https://jev-ai.pro/api
export TYPESAFE_API_KEY=$JEV_AI_API_KEY

That makes a useful development loop: build and test against a local model with no key and no network, then point the same code at hosted Jev through Jev AI’s API when accuracy matters. Re-check thresholds when you switch; the models are not calibrated alike.

Our measurement: 4 Ollama models against hosted Jev

Same items, same question wording, same scoring as our earlier Jev tests. Local models ran one request at a time.

HaluEval judging accuracy

  • Jev91.3%
  • nimble86.7%
  • tev1 (4B)84.3%
  • tev1:0.8b70.3%
  • laya69.0%

BFCL tool routing accuracy

  • Jev93.0%
  • nimble90.7%
  • tev1 (4B)93.0%
  • tev1:0.8b86.8%
  • layan/a

Declines when no tool fits

  • Jev87.5%
  • nimble83.3%
  • tev1 (4B)87.9%
  • tev1:0.8b77.9%
  • layan/a
ModelWhereJudge accuracyAccepts · rejectsJudge ECERouting accuracyRight tool · declinesRouting ECEMedian per verdict
JevTypeSafe · jev-1.13.0, size undisclosedHosted API, from Asia91.3%95.3% · 87.3%0.05693.0%99.5% · 87.5%0.046535 ms
nimbleBespoke Labs · 9B, Qwen3.5-9BOllama, laptop86.7%96.0% · 77.3%0.10690.7%99.5% · 83.3%0.069289 ms
tev1 (4B)Together AI · 4B, Qwen3.5Ollama, laptop84.3%96.0% · 72.7%0.10593.0%99.0% · 87.9%0.023138 ms
tev1:0.8bTogether AI · 0.8B, Qwen3.5Ollama, laptop70.3%90.0% · 50.7%0.17386.8%97.5% · 77.9%0.04939 ms
layaConvai Innovations · 421M, ModernBERT-largeOllama, laptop69.0%90.7% · 47.3%0.184Not runnable: tool descriptions exceed its token budget18 ms

What we found

  • No local model matched hosted Jev as a judge. The closest, nimble, scored 86.7% against Jev’s 91.3% on 300 HaluEval verdicts; the gap is almost all in catching wrong answers, 77.3% rejected against Jev’s 87.3%.
  • Tool routing is closer. tev1 (4B) tied Jev at 93.0% on 440 BFCL decisions with calibration error 0.023, and every model except laya picked the right tool at least 97.5% of the time when one fitted. Declining when no tool fits separates them: Jev 87.5%, tev1 87.9%, nimble 83.3%, tev1:0.8b 77.9%.
  • The small models accept almost everything. tev1:0.8b and laya passed about 90.0% of correct answers but rejected only 50.7% and 47.3% of hallucinated ones, close to a coin flip, and their probabilities were the least calibrated (ECE 0.173 and 0.184).
  • Speed tracks size. Median per judging verdict: laya 18 ms, tev1:0.8b 39 ms, tev1 138 ms, nimble 289 ms on the laptop, against 535 ms for Jev over the network. Our states are long (about 300 tokens for judging, 565 for routing), so these are slower than Ollama’s 91 ms figure for short game states.
  • laya could not run the routing test at all: it caps instructions at about 31 tokens and each option at 48, and BFCL’s tool descriptions are longer. That limit is worth knowing before you point an existing Jev workload at it.

How we tested

  • Judging: 150 HaluEval QA items, each with a correct and a hallucinated answer, so 300 Noul verdicts on “is the candidate answer correct according to the passage”. Identical to our Jev-as-a-Judge test.
  • Routing: 200 BFCL “multiple” requests with two to four tools and one right answer, plus 240 “irrelevance” requests where the right answer is no tool, as one Choice per request. Identical to our Jev agent test.
  • Setup: Ollama 0.40.2, default quantisation of each library tag, an Apple M5 Pro MacBook with 48 GB, one request at a time so latency is per decision. Hosted Jev ran at concurrency 4 from Asia.
  • Limits: two public datasets, one wording, one run each, no prompt tuning for any model. Calibration error is reported but probabilities on small models cluster near the extremes, so read it with the accuracy.

What differs from hosted Jev

Same request shape, different limits.

Ollama decision modelsHosted Jev
Options per Choice or Score2 to 26 (tev1 trained on 2 to 24; laya also caps option text at 48 tokens)2 to 255 for Choice, 2 to 10 for Score
Questions per request1 to 64Many; limited by the context and request size
Model namesnimble, tev1, tev1:0.8b, clef, clef-flash, layajev-latest, jev-1.13.0
Imagesclef and clef-flash only (Ollama 0.35.1+)Text only; see Is Jev multimodal?
AuthenticationNone; any key is acceptedBearer key
Extra response fieldsAdds prompt_eval_cached_count at the top levelmodel, answers, usage
Where it runsYour machine; nothing leaves itHosted API

Two practical traps. Ollama model names contain a colon (tev1:0.8b), which some clients reject as a model ID; ollama cp tev1:0.8b tev1-small makes an alias. And the extra prompt_eval_cached_count field breaks clients that refuse unknown fields; we hit this with OpenClaw’s TypeSafe plugin.

Every decision model on Ollama

Ollama’s library on 2026-10-11. Ollama’s announcement named three; three more have been added since.

ModelFromSizeImagesMeasured hereNote
nimbleBespoke Labs9BNoYesAnnounced with the endpoint on 29 September
tev1Together AI4BNoYesExperimental; 0.8B variant as tev1:0.8b
tev1:0.8bTogether AI0.8BNoYesExperimental
clefCloudflare27BYesNoNeeds Ollama 0.35.1+ for images; see our Clef guide
clef-flashCloudflare9BYesNoLatency-focused Clef; 12 GB download
layaConvai Innovations421MNoYesEncoder; strict token budget per instruction and option

nimble

Bespoke Labs · 9B, Qwen3.5-9B · 9.3 GB download · Apache-2.0

Reads the prompt once per question and scores the answer tokens directly, with no reasoning step. Trained on contrastive pairs that differ by one fact. Up to 64 questions per request.

tev1 (4B)

Together AI · 4B, Qwen3.5 · 4.4 GB download · MIT training code and data builders

An experimental decision model meant for the routing and policy checks Jev does. Trained on 2 to 24 options per question.

tev1:0.8b

Together AI · 0.8B, Qwen3.5 · 796 MB download · MIT training code and data builders

The smallest Ollama decision model in our test, for the tightest latency or memory budgets.

laya

Convai Innovations · 421M, ModernBERT-large · 846 MB download · Apache-2.0

An encoder with a decision head; under 10 ms per question on an M5 Max by Ollama’s figure. Instructions and each option share a fixed token budget, so long option text is rejected.

Other open options that speak the same API: Kev, Jeeves and pplx-decider. The full list is on Can you run Jev locally?

Jev on Ollama FAQ

Is Jev available on Ollama?

No. TypeSafe has not released Jev’s weights, so there is no Jev model to pull. What Ollama added in version 0.35, on 29 September 2026, is Jev’s API: a /v1/systemone endpoint that takes the same state and typed questions, served by open decision models: nimble, tev1, tev1:0.8b, clef, clef-flash and laya as of 11 October.

Which Ollama decision model is closest to Jev?

It depends on the job. As a judge, nimble came closest: 86.7% on 300 HaluEval verdicts against Jev’s 91.3%. For tool routing, tev1 (4B) scored 93.0% on 440 decisions against Jev’s 93.0%. The gap is largest when the right answer is no: rejecting a wrong answer, or declining when no tool fits.

How fast are decision models on Ollama?

laya answered a judging verdict in a 18 ms median on an Apple M5 Pro, one request at a time. Hosted Jev took 535 ms from Asia including the network. Ollama’s own figure is 91 ms per decision for nimble on an M5 Max in a game loop.

Does my Jev code work with Ollama unchanged?

Mostly. Change the base URL to http://localhost:11434 and the model name, and any key is accepted. Keep Choice and Score questions to 26 options or fewer. Clients that validate responses strictly can reject Ollama’s extra prompt_eval_cached_count field; the OpenClaw TypeSafe plugin does, as our OpenClaw page shows.

What hardware do I need?

It depends on the model: nimble is a 9.3 GB download, tev1 (4B) is a 4.4 GB download, tev1:0.8b is a 796 MB download, laya is a 846 MB download. We ran all of them on a 48 GB Apple M5 Pro with room to spare. CPUs work for the smaller models at lower speed.

When should I use hosted Jev instead?

When accuracy on hard or subtle decisions matters more than keeping data on the machine, when you need more than 26 options, or when you do not want to run inference yourself. A common split is local for high-volume easy gates and Jev for the ones that decide money, safety or customers.

About this page

Who. Jev AI runs a hosted endpoint for TypeSafe’s Jev and publishes measurements of it and its alternatives. We are not affiliated with Ollama, Bespoke Labs, Together AI or TypeSafe.

How. Model facts are from Ollama’s blog post and library pages. All accuracy, calibration and latency numbers are our own runs on 2026-10-11 with the scripts and result files in our repository.

When. Published October 11, 2026. Ollama says more decision models and MLX acceleration on Apple silicon are coming; we will rerun when they ship.

Sources

Ollama and the models named here belong to their respective owners and are not affiliated with Jev AI or TypeSafe.