Cygnet
68.6rank #7 · Capability #7
Model guide · Untrained recipe · frozen Gemma 4 12B on vLLM · MIT shim · updated October 7, 2026
Seventh on the board, and nothing about the model is trained.
Cygnet does not train anything. It serves Google’s Gemma 4 12B unchanged on stock vLLM, and a short MIT-licensed shim turns each decision into one masked answer letter whose probability it reads and calibrates with a single temperature. That was enough for first place in JevBench v1.5.6 and #7 of 92 on the re-measured v1.6.0 pool. It is the clearest evidence that a general model, read the right way, can do much of what a trained decision model does.
Cygnet is an independent project by blockbrain. Jev AI does not host it; this page is a guide to what it is and how to run it yourself.
One benchmark measured every system under one method and ranked 92. The full board is on the JevBench results page.
Cygnet
68.6rank #7 · Capability #7
Jev 1.13.0
71.1rank #2 · Capability #3
IntelligenceHow often it picks the right answer, half from sealed decisions
CalibrationWhether 0.8 really means about 80%
SpeedMeasured response time
Open half 58.3 · sealed half 51.3; chance-corrected by request type: choice 69.1, yes/no 48.7, score 46.8. The composite also weighs a fourth axis, cost, which these bars leave out. Self-hosted latency is adjusted by the benchmark to approximate production load.
For each decision the shim sends one chat request with the state, the instructions and the options relabelled A, B, C and so on, and asks for one letter. vLLM restricts the answer position to those letters and returns their log-probabilities; the shim sums every token that decodes to each letter, renormalises over the options and applies a calibration temperature of 3.4. Output is one token per decision. The author credits the one-token readout to NInfer.
The temperature was fitted on 241 items the author generated; the README states that JevBench’s public items were used only to measure, never to fit. On the public set the author measured 203 of 231 correct on both an RTX A6000 and an L40S, with a 50–66ms median on the standard tier.
The repository appeared on 24 September 2026. It is a recipe you run, not a service: there is no hosted Cygnet to call.
JevBench v1.6.0 ranks Cygnet #7 of 92 at 68.6, measured on an RTX PRO 6000 at a 37ms raw median. Speed is 91.8 and calibration 87.0; intelligence is 54.8, with 58.3 on the open half and 51.3 on the sealed half. Jev is #2 at 71.1 with intelligence 62.4 and leads on all three request types. On the Capability ranking Cygnet is seventh at 70.9, ahead of every 4B rebuild documented on this site, and deck-31B at #5 is the only other untrained system above it.
Serve gemma-4-12B-it on vLLM at the pinned revision, start the shim in front of it, and send /v1/systemone requests to the shim. The author’s runs used 48 GB GPUs.
git clone https://github.com/blockbrain-ai/cygnet-recipe
cd cygnet-recipe
# Follow README.md: vLLM 0.30.0 serving google/gemma-4-12B-it at the pinned revision,
# then the shim, which accepts JevBench's /v1/systemone requests.
curl http://localhost:8000/v1/systemone \
-H "Content-Type: application/json" \
-d '{"state": "Chat: Can you cancel my subscription and refund this month?",
"questions": {"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"cancel": null, "refund": null, "both": null, "other": null}}}}'From the project’s own documentation and JevBench v1.6.0.
Cygnet gets remarkably far with no training, which makes it worth measuring on your own cases rather than taking either ranking on trust. Run them on Jev here with no setup and keep the answers; the gap on your decisions is the only one that matters.
A recipe from blockbrain: Google’s gemma-4-12B-it served unchanged on stock vLLM, with a short MIT shim that turns each typed decision into one masked answer letter and reads its probability. Nothing is trained or merged.
No. The only fitted value is one calibration temperature, 3.4, which the author fitted on 241 self-generated items rather than on JevBench’s public set.
#7 of 92 at 68.6 in JevBench v1.6.0, with speed 91.8 and calibration 87.0. It led the previous release, v1.5.6, on the old pool; the full re-measure reversed that and Jev is now #2 at 71.1.
The author measured it on an RTX A6000 and an L40S, both 48 GB, and JevBench measured it on an RTX PRO 6000. Other cards have not been measured.
No. Both use Gemma 4 12B, but Winnow fine-tunes it and Cygnet uses the stock model. They are independent projects by different authors.
Every figure on this page is quoted from a published source and was read on October 7, 2026.
Cygnet and every other product named on this page belongs to its respective owner and is not affiliated with Jev AI. Figures are quoted from the sources above as they read on October 7, 2026 and can change.