Jev AI

Jev for search · project guide and measurement · published October 8, 2026

Jev Search

An open-source metasearch that lets Jev pick where to search and rank what comes back, and what that tells you about Jev as a reranker.

Jev Search, by Search1API, is the most-starred public project built on Jev: a web search that generates no answers. Jev decides which of twelve engines to query, over what time range and with which query string, then scores every result for relevance. This page explains how it works, collects the retrieval benchmarks published for Jev and Clef, and adds a measurement of our own on a labelled dataset, with the data and script published.

521GitHub stars, 62 forks, MIT, active since 17 September 2026
12engines queried concurrently through Search1API
0.608 → 0.708nDCG@10 after Jev reranks BM25 in our SciFact test
$0.001to rerank 20 candidates for one query

How Jev Search works

Three stages, two of them decisions. The application’s own description, condensed.

  1. 01

    Understand

    Jev answers typed questions about the request: which engines fit, which time range, what query string to send. No text is generated; each answer is a choice with probabilities.

  2. 02

    Search

    Search1API runs the chosen query on the chosen engines concurrently, each with a 15-second deadline inside a 30-second request budget. Successful responses are cached for 10 minutes to 6 hours depending on the time window.

  3. 03

    Rank

    Jev scores every result for relevance to the request. Results are merged by URL, ordered by relevance, engine agreement and original rank, streamed as each engine finishes, and the low scorers are grouped apart.

Engines: Google · DuckDuckGo · Yandex · Hacker News · Reddit · GitHub · X · arXiv · YouTube · Wikipedia · IMDb · WeChat. Limits: 10 searches per IP per minute per Cloudflare location by default. Stack: Cloudflare Workers, a KV namespace for caching, Search1API for the searching, and any one of the decision models below for the judging.

The decision models it can use

Jev Search added a model selector on 2 October 2026 and GPT-6 Luna on 7 October. Latency and price are the figures its selector shows, checked October 2026.

ModelByLatency shownPrice shownOn Jev AI
JevTypeSafe524.1 ms$0.042 / M tokens via TypeSafeGuide
ClefCloudflare209.3 ms$0.24 / M tokens on Workers AIGuide
Clef FlashCloudflare38.8 ms$0.09 / M tokens on Workers AIGuide
GPT-6 LunaOpenAINot listed$0.10 / M input tokensGuide

Latency is what the project measured for its own calls; the Jev 1.13.0 guide has the version’s limits and JevBench’s 0.24 s median for the hosted API.

What has been measured about Jev for retrieval

Two public numbers from Cloudflare’s Clef launch, one independent account, and the project’s own caveats.

Retrieval benchmarks: Jev against Clef

Cloudflare, Clef launch article · October 1, 2026

BenchmarkMetricJevClefClef Flash
ToolRetnDCG@1065.2869.1966.43
BRIGHTnDCG@1047.5245.9139.26

ToolRet is tool and API retrieval; BRIGHT is reasoning-heavy retrieval. Clef leads the first, Jev the second, and Clef Flash trails both while being more than ten times faster. These are Cloudflare’s own measurements of all three models.

Read the source ↗ · Jev vs Clef on this site

A search company’s own trial

Parallel, “Testing Jev” · September 2026

Reranking
NDCG@10 of 0.7 ordering candidate documents by relevance to a query, “comparable to internal system”
Topic classification
Their internal models did better; choosing from a large label set is named as a weakness
Query freshness
Internal models did better; the task is suspected to be out of Jev’s training distribution
Latency and cost
Competitive latency against their larger models; materially higher cost per document than their own classifiers at their scale

The one task where Jev matched a search company’s in-house system is the one Jev Search uses it for: ranking. The authors note their comparison is against dedicated classifiers, not general LLMs.

Read the source ↗

Our own test: Jev reranking BM25 on SciFact

Jev Search scores each result after the engines return. We did the same thing on a dataset with answer keys, on 2026-10-08, and published the script.

DatasetBEIR SciFact, test split: 60 scientific claims, every 5th of 300 test queries with relevance labels, over a corpus of 5,183 abstracts with human relevance labels
First stagePlain BM25 over title and abstract, top 20 per claim; 57 of the 71 labelled abstracts were among the candidates
RerankOne Score question per candidate, “How relevant is this abstract to the scientific claim?”, four described levels; candidates reordered by the returned score, BM25 order as tie-break
Judgments1,200 on jev-latest (jev-1.13.0) through TypeSafe’s API at list price, 2026-10-08
nDCG@10BM25 0.608 → Jev 0.708; a perfect reordering of the same candidates would reach 0.791
Recall@10 · MRR@10BM25 0.767 · 0.563 → Jev 0.788 · 0.686
Per query19 improved, 5 got worse, 36 unchanged on nDCG@10
Latency323 ms median, 406 ms p95 per candidate; Client-observed round trip per candidate from a laptop in Asia at concurrency 4.
Tokens and cost771 input tokens per judgment; $0.039 for the run, $0.001 per query of 20 candidates, $0.032 per thousand judgments

What each score level contained

For every level Jev assigned, how many candidates landed there and what share of them were labelled relevant. A useful reranker puts the relevant abstracts at the top levels and almost none at the bottom.

ScoreLevelCandidatesShare labelled relevant
0Unrelated to the claim: a different subject entirely.7890.1%
1Same subject, but it does not address whether the claim is true.2633.0%
2Addresses the claim partially or indirectly, or only one part of it.9316.1%
3Directly reports evidence that supports or refutes the claim.5560.0%

What we found

  • Reranking the same 20 candidates moved nDCG@10 by +10 points, from 0.608 to 0.708, against a ceiling of 0.791 for a perfect reorder. Recall@10 went from 0.767 to 0.788.
  • 19 of 60 queries improved and 5 got worse; the rest already had every relevant abstract in the top ten or none among the candidates.
  • Candidates Jev scored at the top level were relevant 60% of the time, against 0% at the bottom level, which is the separation a threshold can use.
  • The cost of a query, $0.001 for 20 judgments, is the figure to compare with a cross-encoder you host yourself; the latency, 323 ms per candidate from Asia, is why Jev Search streams results and why Clef Flash exists as an option.

Largest moves of a relevant abstract

  • BM25 #16 → Jev #3Epidemiological disease burden from noncommunicable diseases is more prevalent in low economic settings.Global, regional, and national comparative risk assessment of 79 behavioural, environmental and occupational, and metabolic risks or clusters of risks, 1990–2015: a systematic analysis for the Global Burden of Disease Study 2015
  • BM25 #10 → Jev #1Bone marrow cells contribute to adult macrophage compartments.Tissue-resident macrophages self-maintain locally throughout adult life with minimal contribution from circulating monocytes.
  • BM25 #12 → Jev #3Women with a higher birth weight are more likely to develop breast cancer later in life.Intrauterine factors and risk of breast cancer: a systematic review and meta-analysis of current evidence.
  • Pushed down: BM25 #2 → Jev #8Cold exposure increases BAT recruitment.Cold Exposure Promotes Atherosclerotic Plaque Growth and Instability via UCP1-Dependent Lipolysis

Limits of this test. One dataset of scientific claims, one first-stage retriever, one question wording, one run of 60 queries. BM25 here is a plain implementation without stemming, so its baseline is lower than a tuned one; the oracle row shows how much of the gap a reranker could close at all. Latency includes a long network path. The labels are the dataset’s own; nothing was relabelled or filtered after the run. Script, sampling rule and full output are in the repository file named above.

Rerank with the Jev AI API

The request our test sent for each candidate. Levels describe situations, not degrees, which is what TypeSafe’s Score guidance asks for.

curl --fail-with-body https://jev-ai.pro/api/v1/systemone \
  -H "Authorization: Bearer $JEV_AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": {
      "claim": "Epidemiological disease burden from noncommunicable diseases is more prevalent in low economic settings.",
      "abstract_title": "…",
      "abstract": "…"
    },
    "questions": {
      "relevance": {
        "type": "score",
        "instructions": "How relevant is this abstract to the scientific claim?",
        "criteria": [
                  "Unrelated to the claim: a different subject entirely.",
                  "Same subject, but it does not address whether the claim is true.",
                  "Addresses the claim partially or indirectly, or only one part of it.",
                  "Directly reports evidence that supports or refutes the claim."
        ]
      }
    }
  }'
  1. 01

    Retrieve first, judge second

    Let BM25, an embedding index or a search API produce candidates. Jev reads text; it does not search.

  2. 02

    One candidate per request, or many questions per state

    We judged candidates one at a time. Jev also answers many questions about one state in parallel, so a page of results can be scored in one call if the state stays well under the 32k-token state limit.

  3. 03

    Reorder on the score, keep the probabilities

    Sort by the returned score with the original rank as the tie-break. The probability spread tells you which results to drop or group apart, as Jev Search does with its low scorers.

  4. 04

    Choose the model by latency budget

    Jev for accuracy on reasoning-heavy queries, Clef Flash when twenty judgments must finish inside a second. Both run from the same request on Jev AI; see the Clef guide.

Related tools on this site: Live Web Context runs a web search for a yes/no question and lets Jev weigh the evidence; jevgrep does the same for code; RAG evaluation scores retrieved passages before generation.

Jev Search FAQ

What is Jev Search?

An open-source metasearch application by Search1API (SuperAgents Lab) that uses TypeSafe’s Jev as its decision layer: Jev chooses which of twelve engines to query and over what time range, then scores every result for relevance. It generates no text, runs on Cloudflare Workers, is MIT-licensed and has a public demo at jev.s1.dev. It is not a TypeSafe product.

Is Jev a search engine?

No. Jev answers typed questions about text with probabilities. Jev Search uses those answers to drive ordinary search engines and to rank what they return. The searching is done by Search1API; the judging is done by Jev.

Can Jev rerank search results?

Yes, and it is one of the better-measured uses of the model. Cloudflare’s October 2026 comparison scores Jev at 65.28 nDCG@10 on ToolRet and 47.52 on BRIGHT. Our own test on 60 SciFact claims moved BM25’s nDCG@10 from 0.608 to 0.708 by asking one Score question per candidate.

How much does reranking with Jev cost?

Jev bills input tokens only, at $0.042 per million. In our test a candidate judgment used about 771 tokens, so reranking 20 candidates cost $0.001 per query, or $0.032 per thousand judgments. Search1API’s own guide warns that at very large per-document volumes a dedicated classifier can be cheaper.

Is Jev or Clef better for search?

Split, by Cloudflare’s own numbers: Clef leads on ToolRet (69.19 against 65.28 nDCG@10) and Jev leads on BRIGHT (47.52 against 45.91). Clef Flash is far faster, 38.8 ms against Jev’s 524 ms in Jev Search’s listing, which matters when you score twenty results per query. Both run on Jev AI.

How do I use Jev for search ranking myself?

Send the query and one candidate per request with a Score question whose levels describe relevance in words, read the score and probabilities, and reorder. The request example on this page is the one our test used; the Live Web Context tool on this site runs a related loop against a web search for yes/no questions.

Does Jev Search send my queries to TypeSafe?

The application sends the request text to its configured decision provider to answer the typed questions, and the chosen query to Search1API. Its README states it does not record search queries or result clicks in its own analytics; check its privacy disclosure and the provider’s terms before deploying it with user data.

About this page

How it was made, so you can judge it.

Who. Jev AI operates an independent playground and API for Jev. We are not affiliated with TypeSafe, Search1API, Cloudflare or Parallel. We sell access to the model this page is about, which is why the third-party figures are quoted with their sources and our own test is published in full.

How. Project facts are from the Jev Search repository, its README and its demo as read on October 8, 2026. Benchmark figures are quoted from Cloudflare’s launch article and Parallel’s post. Our measurement used TypeSafe’s public API at list price and the public BEIR SciFact test split.

When. Published October 8, 2026. Jev Search changes weekly and jev-latest currently resolves to jev-1.13.0; the page will say so when either moves on.

Sources

Jev is developed by TypeSafe. Jev Search, Search1API, Clef, GPT-6 Luna and the engines named above belong to their respective owners and are not affiliated with Jev AI.