Winnow-12B
68.9rank #6 · Capability #6
Model guide · Gemma 4 12B decision fine-tune · GGUF · decisions, chat and vision · updated October 7, 2026
One local model that decides, chats and looks at a screenshot.
Winnow-12B fine-tunes Google’s Gemma 4 12B for Jev-style typed decisions and serves them from a llama.cpp-based server that also answers ordinary chat and takes images through an optional projector. Its author tested the 8-bit build with a 64K context and vision on a 16 GB RTX 5070 Ti. JevBench v1.6.0 ranks it #6 of 92, the strongest system on the board that also chats.
Winnow-12B is an independent project by EldanRing. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.
One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.
Winnow-12B
68.9rank #6 · Capability #6
Jev 1.13.0
71.1rank #2 · Capability #3
IntelligenceHow often it picks the right answer, half from sealed decisions
CalibrationWhether 0.8 really means about 80%
SpeedMeasured response time
Open half 59.8 · sealed half 59.2; chance-corrected by request type: choice 70.1, yes/no 51.5, score 57.0. The composite also weighs a fourth axis, cost, which these bars leave out. Self-hosted latency is adjusted by the benchmark to approximate production load.
Winnow-12B is a LoRA fine-tune of google/gemma-4-12B-it merged into the weights and released as two GGUF files: an 8-bit Q8_0 build recommended for 16 GB GPUs and a 16-bit BF16 build. It runs in the author’s llama.cpp-based inference server; no conversion step or separate adapter is needed.
For decisions the server fills in the shared state once, forks one branch per question and reads the logits of the answer tokens without generating text, the same idea as Jev’s typed questions. It supports noul, choice and score. The same loaded model serves chat on /v1/chat/completions, and the optional projector adds image input. It states that it is not affiliated with or endorsed by TypeSafe or Google.
The training data is private. The model card describes curated typed-decision examples covering routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase and answerability, followed by a refinement pass.
JevBench v1.6.0 ranks Winnow-12B Q8 #6 of 92 at 68.9, measured on an RTX 6000 at a 0.12s raw median per decision. Intelligence is 59.5, with 59.8 on the open half and 59.2 on the sealed half; calibration is 83.0 and speed 86.7. Jev is #2 at 71.1 with intelligence 62.4 and calibration 90.6. Winnow was second, above Jev, in the previous release on the old pool, and holds sixth place under all three of the benchmark’s orders in v1.6.0.
Its model card also reports a second suite, Kev-v9 clean, with 1,046 items, where the author scores Jev at 87.00% and Winnow Q8 at 81.55%, and a tie at 85.71% on the 231-item public JevBench subset. Those are the author’s measurements, not JevBench’s.
Download the Q8_0 GGUF and, if you want image input, the projector; then start the author’s inference server. One process serves decisions and chat.
# Weights: huggingface.co/EldanRing/Winnow-12B (Q8_0 ~12.7 GB, BF16 ~23.8 GB, projector 175 MB)
git clone https://github.com/EldanRing/winnow-inference
cd winnow-inference
# Follow README.md to build the llama.cpp-based server and point it at the GGUF.
curl http://localhost:8080/v1/systemone \
-H "Content-Type: application/json" \
-d '{"state": "Refund request: item returned unopened 40 days after delivery.",
"questions": {"eligible": {"type": "noul", "instructions": "Is this within a 30-day return window?"}}}'From the project’s own documentation and JevBench v1.6.0.
Winnow does more on one GPU than anything else on this site, so the useful comparison is on your own cases and your own constraints. Run them on Jev here with no setup and keep the answers; if Winnow agrees on the decisions you care about and you can run it, a local model that also chats and reads images may still be the better fit.
An Apache-2.0 fine-tune of Google’s Gemma 4 12B by EldanRing that answers Jev-style noul, choice and score questions from a llama.cpp-based server, and also serves chat and image input from the same loaded model.
The author tested the 8-bit Q8_0 build, about 12.7 GB, with full offload on a 16 GB RTX 5070 Ti at a 64K context. The BF16 build, about 23.8 GB, needs more memory or offloading.
Not in JevBench v1.6.0: it is #6 at 68.9 against Jev’s #2 at 71.1, with intelligence 59.5 against 62.4. It was ahead of Jev in the previous release on the old pool.
Yes, with the optional 175 MB vision projector loaded alongside the GGUF. Jev is text only.
Its server exposes /v1/systemone for typed decisions with noul, choice and score questions, and /v1/chat/completions for chat. Test your own requests against it before switching an integration.
Every figure on this page is quoted from a published source and was read on October 7, 2026.
Winnow-12B and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.