Jev AI

Model guide · Gemma 4 12B decision fine-tune · GGUF · decisions, chat and vision · updated October 7, 2026

Winnow-12B

One local model that decides, chats and looks at a screenshot.

Winnow-12B fine-tunes Google’s Gemma 4 12B for Jev-style typed decisions and serves them from a llama.cpp-based server that also answers ordinary chat and takes images through an optional projector. Its author tested the 8-bit build with a 64K context and vision on a 16 GB RTX 5070 Ti. JevBench v1.6.0 ranks it #6 of 92, the strongest system on the board that also chats.

Winnow-12B is an independent project by EldanRing. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.

Developer
EldanRing (independent)
Base model
google/gemma-4-12B-it, LoRA merged
Builds
Q8_0 about 12.7 GB · BF16 about 23.8 GB
Licence
Apache-2.0, with Gemma 4 terms for derivatives
Inputs
Text, plus images with a 175 MB vision projector
Context
65,536 positions tested with vision on 16 GB
Endpoints
/v1/systemone and /v1/chat/completions
Tested GPU
RTX 5070 Ti, 16 GB, Q8 build

Winnow-12B on JevBench v1.6.0

One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.

Winnow-12B

68.9rank #6 · Capability #6

Self-hosted, RTX 6000 · 0.12s raw, 0.38s adjusted

Jev 1.13.0

71.1rank #2 · Capability #3

Hosted API · 0.24s median / 0.30s p95

IntelligenceHow often it picks the right answer, half from sealed decisions

Winnow-12B59.5
Jev62.4

CalibrationWhether 0.8 really means about 80%

Winnow-12B83.0
Jev90.6

SpeedMeasured response time

Winnow-12B86.7
Jev91.5

Open half 59.8 · sealed half 59.2; chance-corrected by request type: choice 70.1, yes/no 51.5, score 57.0. The composite also weighs a fourth axis, cost, which these bars leave out. Self-hosted latency is adjusted by the benchmark to approximate production load.

What Winnow-12B is

Winnow-12B is a LoRA fine-tune of google/gemma-4-12B-it merged into the weights and released as two GGUF files: an 8-bit Q8_0 build recommended for 16 GB GPUs and a 16-bit BF16 build. It runs in the author’s llama.cpp-based inference server; no conversion step or separate adapter is needed.

For decisions the server fills in the shared state once, forks one branch per question and reads the logits of the answer tokens without generating text, the same idea as Jev’s typed questions. It supports noul, choice and score. The same loaded model serves chat on /v1/chat/completions, and the optional projector adds image input. It states that it is not affiliated with or endorsed by TypeSafe or Google.

The training data is private. The model card describes curated typed-decision examples covering routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase and answerability, followed by a refinement pass.

Where it stands on JevBench

JevBench v1.6.0 ranks Winnow-12B Q8 #6 of 92 at 68.9, measured on an RTX 6000 at a 0.12s raw median per decision. Intelligence is 59.5, with 59.8 on the open half and 59.2 on the sealed half; calibration is 83.0 and speed 86.7. Jev is #2 at 71.1 with intelligence 62.4 and calibration 90.6. Winnow was second, above Jev, in the previous release on the old pool, and holds sixth place under all three of the benchmark’s orders in v1.6.0.

Its model card also reports a second suite, Kev-v9 clean, with 1,046 items, where the author scores Jev at 87.00% and Winnow Q8 at 81.55%, and a tie at 85.71% on the 231-item public JevBench subset. Those are the author’s measurements, not JevBench’s.

Run Winnow-12B

Download the Q8_0 GGUF and, if you want image input, the projector; then start the author’s inference server. One process serves decisions and chat.

# Weights: huggingface.co/EldanRing/Winnow-12B (Q8_0 ~12.7 GB, BF16 ~23.8 GB, projector 175 MB)
git clone https://github.com/EldanRing/winnow-inference
cd winnow-inference
# Follow README.md to build the llama.cpp-based server and point it at the GGUF.
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"state": "Refund request: item returned unopened 40 days after delivery.",
       "questions": {"eligible": {"type": "noul", "instructions": "Is this within a 30-day return window?"}}}'
  • On the author’s RTX 5070 Ti, a four-question request at near-full 64K context took 25 seconds cold and a 143ms median once the context was cached, peaking at 15.01 GiB of GPU memory including the desktop.
  • The BF16 build needs more GPU memory or offloading.
  • Confidence is entropy-based and the model card says it is not a guaranteed probability of correctness; validate thresholds on your own labelled data.

Limits to plan around

From the project’s own documentation and JevBench v1.6.0.

  • No fitted calibration map: JevBench v1.6.0 scores its calibration at 83.0 against Jev’s 90.6.
  • Private training data, described by category only.
  • A 12.7 GB model file and a GPU to keep serving; chat and decision contexts share the same process.

Measure it against Jev on your own cases

Winnow does more on one GPU than anything else on this site, so the useful comparison is on your own cases and your own constraints. Run them on Jev here with no setup and keep the answers; if Winnow agrees on the decisions you care about and you can run it, a local model that also chats and reads images may still be the better fit.

Winnow-12B FAQ

What is Winnow-12B?

An Apache-2.0 fine-tune of Google’s Gemma 4 12B by EldanRing that answers Jev-style noul, choice and score questions from a llama.cpp-based server, and also serves chat and image input from the same loaded model.

What GPU do I need for Winnow-12B?

The author tested the 8-bit Q8_0 build, about 12.7 GB, with full offload on a 16 GB RTX 5070 Ti at a 64K context. The BF16 build, about 23.8 GB, needs more memory or offloading.

Is Winnow as accurate as Jev?

Not in JevBench v1.6.0: it is #6 at 68.9 against Jev’s #2 at 71.1, with intelligence 59.5 against 62.4. It was ahead of Jev in the previous release on the old pool.

Can Winnow read images?

Yes, with the optional 175 MB vision projector loaded alongside the GGUF. Jev is text only.

Does Winnow speak the Jev API?

Its server exposes /v1/systemone for typed decisions with noul, choice and score questions, and /v1/chat/completions for chat. Test your own requests against it before switching an integration.

More model guides

Every figure on this page is quoted from a published source and was read on October 7, 2026.

Sources

Winnow-12B and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.