Jev for search · project guide and measurement · published October 8, 2026
Jev Search
An open-source metasearch that lets Jev pick where to search and rank what comes back, and what that tells you about Jev as a reranker.
Jev Search, by Search1API, is the most-starred public project built on Jev: a web search that generates no answers. Jev decides which of twelve engines to query, over what time range and with which query string, then scores every result for relevance. This page explains how it works, collects the retrieval benchmarks published for Jev and Clef, and adds a measurement of our own on a labelled dataset, with the data and script published.
How Jev Search works
Three stages, two of them decisions. The application’s own description, condensed.
- 01
Understand
Jev answers typed questions about the request: which engines fit, which time range, what query string to send. No text is generated; each answer is a choice with probabilities.
- 02
Search
Search1API runs the chosen query on the chosen engines concurrently, each with a 15-second deadline inside a 30-second request budget. Successful responses are cached for 10 minutes to 6 hours depending on the time window.
- 03
Rank
Jev scores every result for relevance to the request. Results are merged by URL, ordered by relevance, engine agreement and original rank, streamed as each engine finishes, and the low scorers are grouped apart.
Engines: Google · DuckDuckGo · Yandex · Hacker News · Reddit · GitHub · X · arXiv · YouTube · Wikipedia · IMDb · WeChat. Limits: 10 searches per IP per minute per Cloudflare location by default. Stack: Cloudflare Workers, a KV namespace for caching, Search1API for the searching, and any one of the decision models below for the judging.
The decision models it can use
Jev Search added a model selector on 2 October 2026 and GPT-6 Luna on 7 October. Latency and price are the figures its selector shows, checked October 2026.
| Model | By | Latency shown | Price shown | On Jev AI |
|---|---|---|---|---|
| Jev | TypeSafe | 524.1 ms | $0.042 / M tokens via TypeSafe | Guide |
| Clef | Cloudflare | 209.3 ms | $0.24 / M tokens on Workers AI | Guide |
| Clef Flash | Cloudflare | 38.8 ms | $0.09 / M tokens on Workers AI | Guide |
| GPT-6 Luna | OpenAI | Not listed | $0.10 / M input tokens | Guide |
Latency is what the project measured for its own calls; the Jev 1.13.0 guide has the version’s limits and JevBench’s 0.24 s median for the hosted API.
What has been measured about Jev for retrieval
Two public numbers from Cloudflare’s Clef launch, one independent account, and the project’s own caveats.
Retrieval benchmarks: Jev against Clef
Cloudflare, Clef launch article · October 1, 2026
| Benchmark | Metric | Jev | Clef | Clef Flash |
|---|---|---|---|---|
| ToolRet | nDCG@10 | 65.28 | 69.19 | 66.43 |
| BRIGHT | nDCG@10 | 47.52 | 45.91 | 39.26 |
ToolRet is tool and API retrieval; BRIGHT is reasoning-heavy retrieval. Clef leads the first, Jev the second, and Clef Flash trails both while being more than ten times faster. These are Cloudflare’s own measurements of all three models.
A search company’s own trial
Parallel, “Testing Jev” · September 2026
- Reranking
- NDCG@10 of 0.7 ordering candidate documents by relevance to a query, “comparable to internal system”
- Topic classification
- Their internal models did better; choosing from a large label set is named as a weakness
- Query freshness
- Internal models did better; the task is suspected to be out of Jev’s training distribution
- Latency and cost
- Competitive latency against their larger models; materially higher cost per document than their own classifiers at their scale
The one task where Jev matched a search company’s in-house system is the one Jev Search uses it for: ranking. The authors note their comparison is against dedicated classifiers, not general LLMs.
Our own test: Jev reranking BM25 on SciFact
Jev Search scores each result after the engines return. We did the same thing on a dataset with answer keys, on 2026-10-08, and published the script.
| Dataset | BEIR SciFact, test split: 60 scientific claims, every 5th of 300 test queries with relevance labels, over a corpus of 5,183 abstracts with human relevance labels |
|---|---|
| First stage | Plain BM25 over title and abstract, top 20 per claim; 57 of the 71 labelled abstracts were among the candidates |
| Rerank | One Score question per candidate, “How relevant is this abstract to the scientific claim?”, four described levels; candidates reordered by the returned score, BM25 order as tie-break |
| Judgments | 1,200 on jev-latest (jev-1.13.0) through TypeSafe’s API at list price, 2026-10-08 |
| nDCG@10 | BM25 0.608 → Jev 0.708; a perfect reordering of the same candidates would reach 0.791 |
| Recall@10 · MRR@10 | BM25 0.767 · 0.563 → Jev 0.788 · 0.686 |
| Per query | 19 improved, 5 got worse, 36 unchanged on nDCG@10 |
| Latency | 323 ms median, 406 ms p95 per candidate; Client-observed round trip per candidate from a laptop in Asia at concurrency 4. |
| Tokens and cost | 771 input tokens per judgment; $0.039 for the run, $0.001 per query of 20 candidates, $0.032 per thousand judgments |
What each score level contained
For every level Jev assigned, how many candidates landed there and what share of them were labelled relevant. A useful reranker puts the relevant abstracts at the top levels and almost none at the bottom.
| Score | Level | Candidates | Share labelled relevant |
|---|---|---|---|
| 0 | Unrelated to the claim: a different subject entirely. | 789 | 0.1% |
| 1 | Same subject, but it does not address whether the claim is true. | 263 | 3.0% |
| 2 | Addresses the claim partially or indirectly, or only one part of it. | 93 | 16.1% |
| 3 | Directly reports evidence that supports or refutes the claim. | 55 | 60.0% |
What we found
- Reranking the same 20 candidates moved nDCG@10 by +10 points, from 0.608 to 0.708, against a ceiling of 0.791 for a perfect reorder. Recall@10 went from 0.767 to 0.788.
- 19 of 60 queries improved and 5 got worse; the rest already had every relevant abstract in the top ten or none among the candidates.
- Candidates Jev scored at the top level were relevant 60% of the time, against 0% at the bottom level, which is the separation a threshold can use.
- The cost of a query, $0.001 for 20 judgments, is the figure to compare with a cross-encoder you host yourself; the latency, 323 ms per candidate from Asia, is why Jev Search streams results and why Clef Flash exists as an option.
Largest moves of a relevant abstract
- BM25 #16 → Jev #3Epidemiological disease burden from noncommunicable diseases is more prevalent in low economic settings.Global, regional, and national comparative risk assessment of 79 behavioural, environmental and occupational, and metabolic risks or clusters of risks, 1990–2015: a systematic analysis for the Global Burden of Disease Study 2015
- BM25 #10 → Jev #1Bone marrow cells contribute to adult macrophage compartments.Tissue-resident macrophages self-maintain locally throughout adult life with minimal contribution from circulating monocytes.
- BM25 #12 → Jev #3Women with a higher birth weight are more likely to develop breast cancer later in life.Intrauterine factors and risk of breast cancer: a systematic review and meta-analysis of current evidence.
- Pushed down: BM25 #2 → Jev #8Cold exposure increases BAT recruitment.Cold Exposure Promotes Atherosclerotic Plaque Growth and Instability via UCP1-Dependent Lipolysis
Limits of this test. One dataset of scientific claims, one first-stage retriever, one question wording, one run of 60 queries. BM25 here is a plain implementation without stemming, so its baseline is lower than a tuned one; the oracle row shows how much of the gap a reranker could close at all. Latency includes a long network path. The labels are the dataset’s own; nothing was relabelled or filtered after the run. Script, sampling rule and full output are in the repository file named above.
Rerank with the Jev AI API
The request our test sent for each candidate. Levels describe situations, not degrees, which is what TypeSafe’s Score guidance asks for.
curl --fail-with-body https://jev-ai.pro/api/v1/systemone \
-H "Authorization: Bearer $JEV_AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": {
"claim": "Epidemiological disease burden from noncommunicable diseases is more prevalent in low economic settings.",
"abstract_title": "…",
"abstract": "…"
},
"questions": {
"relevance": {
"type": "score",
"instructions": "How relevant is this abstract to the scientific claim?",
"criteria": [
"Unrelated to the claim: a different subject entirely.",
"Same subject, but it does not address whether the claim is true.",
"Addresses the claim partially or indirectly, or only one part of it.",
"Directly reports evidence that supports or refutes the claim."
]
}
}
}'- 01
Retrieve first, judge second
Let BM25, an embedding index or a search API produce candidates. Jev reads text; it does not search.
- 02
One candidate per request, or many questions per state
We judged candidates one at a time. Jev also answers many questions about one state in parallel, so a page of results can be scored in one call if the state stays well under the 32k-token state limit.
- 03
Reorder on the score, keep the probabilities
Sort by the returned score with the original rank as the tie-break. The probability spread tells you which results to drop or group apart, as Jev Search does with its low scorers.
- 04
Choose the model by latency budget
Jev for accuracy on reasoning-heavy queries, Clef Flash when twenty judgments must finish inside a second. Both run from the same request on Jev AI; see the Clef guide.
Related tools on this site: Live Web Context runs a web search for a yes/no question and lets Jev weigh the evidence; jevgrep does the same for code; RAG evaluation scores retrieved passages before generation.
Jev Search FAQ
What is Jev Search?
An open-source metasearch application by Search1API (SuperAgents Lab) that uses TypeSafe’s Jev as its decision layer: Jev chooses which of twelve engines to query and over what time range, then scores every result for relevance. It generates no text, runs on Cloudflare Workers, is MIT-licensed and has a public demo at jev.s1.dev. It is not a TypeSafe product.
Is Jev a search engine?
No. Jev answers typed questions about text with probabilities. Jev Search uses those answers to drive ordinary search engines and to rank what they return. The searching is done by Search1API; the judging is done by Jev.
Can Jev rerank search results?
Yes, and it is one of the better-measured uses of the model. Cloudflare’s October 2026 comparison scores Jev at 65.28 nDCG@10 on ToolRet and 47.52 on BRIGHT. Our own test on 60 SciFact claims moved BM25’s nDCG@10 from 0.608 to 0.708 by asking one Score question per candidate.
How much does reranking with Jev cost?
Jev bills input tokens only, at $0.042 per million. In our test a candidate judgment used about 771 tokens, so reranking 20 candidates cost $0.001 per query, or $0.032 per thousand judgments. Search1API’s own guide warns that at very large per-document volumes a dedicated classifier can be cheaper.
Is Jev or Clef better for search?
Split, by Cloudflare’s own numbers: Clef leads on ToolRet (69.19 against 65.28 nDCG@10) and Jev leads on BRIGHT (47.52 against 45.91). Clef Flash is far faster, 38.8 ms against Jev’s 524 ms in Jev Search’s listing, which matters when you score twenty results per query. Both run on Jev AI.
How do I use Jev for search ranking myself?
Send the query and one candidate per request with a Score question whose levels describe relevance in words, read the score and probabilities, and reorder. The request example on this page is the one our test used; the Live Web Context tool on this site runs a related loop against a web search for yes/no questions.
Does Jev Search send my queries to TypeSafe?
The application sends the request text to its configured decision provider to answer the typed questions, and the chosen query to Search1API. Its README states it does not record search queries or result clicks in its own analytics; check its privacy disclosure and the provider’s terms before deploying it with user data.
About this page
How it was made, so you can judge it.
Who. Jev AI operates an independent playground and API for Jev. We are not affiliated with TypeSafe, Search1API, Cloudflare or Parallel. We sell access to the model this page is about, which is why the third-party figures are quoted with their sources and our own test is published in full.
How. Project facts are from the Jev Search repository, its README and its demo as read on October 8, 2026. Benchmark figures are quoted from Cloudflare’s launch article and Parallel’s post. Our measurement used TypeSafe’s public API at list price and the public BEIR SciFact test split.
When. Published October 8, 2026. Jev Search changes weekly and jev-latest currently resolves to jev-1.13.0; the page will say so when either moves on.
Sources
Jev is developed by TypeSafe. Jev Search, Search1API, Clef, GPT-6 Luna and the engines named above belong to their respective owners and are not affiliated with Jev AI.