Open-Source Jeff Matches Jev's Published 83% With a 2B Model Trained on One Home GPU
TL;DR
Jeff is a family of three small open-weight decision models that speak the same request format as TypeSafe's hosted Jev. You describe a situation, list options in plain words, and get back a calibrated probability per option from one forward pass. The 2B model scores 83.1% on a five-benchmark panel where Jev's published figure is 83.0%, and the 0.8B model decides in about 22 ms on an RTX PRO 6000 and 28 ms on an M4 Max. It was trained on one workstation GPU in about two hours, with synthetic data written by an open model on two DGX Sparks, and it hit 353 points on Hacker News within a day. The catch: on anything that needs multi-step reasoning, it is far behind.
What shipped
The Show HN post went up on September 28, the same day the GitHub repo and the model weights appeared on Hugging Face. There are three checkpoints:
- Jeff-Qwen3.5-0.8B: 0.8B parameters, 1.7 GB at 16-bit, fine-tuned from Qwen3.5-0.8B.
- Jeff-Qwen3.5-2B: 2B parameters, 4.2 GB.
- Jeff-Gemma4-E2B: 2B effective (4.6B stored), 9.3 GB, fine-tuned from Gemma 4 E2B.
Code is MIT, weights are Apache 2.0. The server exposes a /v1/systemone endpoint that accepts Jev-style requests: a state string or object plus named questions of three types, choice (pick one of up to 255 options), noul (a yes/no probability), and score (a point on a scale you describe). Several questions in one request are answered together. On Apple silicon it runs on MLX; elsewhere it runs on PyTorch, GPU or CPU.
The README is explicit that this is an independent project, not affiliated with or endorsed by TypeSafe. The training code started as a fork of AutoJev, an open MIT recipe that fine-tunes a 27B Qwen model to return Jev-style decisions. Jeff kept AutoJev's core design and pushed it down to models small enough to run on a laptop.
This is the second open Jev alternative AI Bacon has covered, after ConvAI's ModernBERT-based Laya on September 20. The two take different routes. Laya is a 421M encoder with a 512 to 1,024 token context. Jeff is a pair of decoder LLMs with a trained answer readout, and it copies Jev's wire format, so existing Jev client code can point at it.
How a decision model works in one forward pass
Jeff never writes text. The prompt lays out the state and the lettered options, the model runs once, and a trained readout takes the logits for the option letters and turns them into a probability distribution. Training is full-weight fine-tuning for one epoch, batches of 256, with cross-entropy over the option letters. After training, the author fits a single temperature so the probabilities are calibrated, meaning a 0.8 should be right about 80% of the time.
Think of it as the difference between an essay exam and a multiple-choice sheet where you shade in how sure you are for each bubble. A chat model writes the essay and you parse it; Jeff just shades the bubbles, and temperature fitting is a teacher going back afterwards to correct for students who shade every bubble at 100%.
The model card reports expected calibration error of 0.049 for the 0.8B and 0.028 for the 2B, against roughly 0.06 for Jev (an average of Jev's per-benchmark figures, as the card notes).
The home lab pipeline
The build story is the reason this took off on Hacker News. In the author's words, "everything ran at home": one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write synthetic training data, a MacBook for testing, "all monitored from my phone over Tailscale."
The README states that no closed-model output went into the training data; a closed model was used only to spot-check a sample of the synthetic data for quality. Checkpoints were selected on a development set, never on the benchmark panel, and the data mix includes a leak filter. About half of each training family follows the panel's layout conventions, formats only, and the author says no panel item was trained on. The training data itself is not released because some sources are share-alike, but every source and its license is listed.
The benchmarks: great at classifying, weak at thinking
The panel is 4,599 questions drawn from five public benchmarks: BBH, Financial PhraseBank, JudgeBench, RAGTruth, and WinoGrande. Overall, Jeff-2B scores 83.1%, Jeff-Gemma4-E2B 81.6%, and Jeff-0.8B 79.1%. Jev's published figure is 83.0% and AutoJev-27B's is 84.9%.
That overall tie hides a split. Jeff wins big on classification and grounding, and loses big on reasoning:
On JevBench's public hard tier, scored separately over 105 items, Jeff-2B gets 53.3% and Jeff-0.8B 47.6%, against Jev's 73.3%. The author does not hide any of this. From the HN post: "On multi-step reasoning it's behind... That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one."
Read the footnote before you quote the tie
Every Jev and AutoJev number in Jeff's table is a published figure, measured on a different sample of the same benchmarks. Jeff's numbers are measured by the author. So "83.1% vs 83.0%" is a cross-study comparison, not a head-to-head run, and a 0.1-point gap across different samples is noise. AutoJev's own README reports slightly different overall figures again (84.6% for AutoJev-27B and 82.8% for Jev), which shows how much these numbers move with the sample. The honest reading: Jeff lands in the same band as Jev on this panel, not above it.
Speed is the actual pitch
Median decision time over 200 benchmark questions of about 200 input tokens each, raw text to probabilities:
The Jev comparison is not apples to apples, and the README says so: Jev's 114 to 212 ms figures come from published Doom runs over its API, network included. But that is the practical comparison. If your code needs a judgment call inside a game loop, a voice UI, or a request path, a local 22 ms answer with no network hop and no per-call bill is a different product from a hosted one, whatever the benchmark says. CPU-only is a different story: 463 ms for the 0.8B on 32 threads is slower than calling Jev.
The fun part: Doom, Frogger, Pac-Man
As a zero-shot test with no game data in training, the author had Jeff play three games for 20 episodes each (seed 1234). Each turn, code describes the situation and legal moves in words, and the options state consequences ("you would be hit by a car and lose a life") but never which move is right.
- Doom (ViZDoom, harness adapted from jev-plays-doom): Jeff-0.8B averaged 6.55 kills per episode, level with a hand-coded rule bot and with Jev's published run. Jev's run was given an explicit aiming rule and scored -0.60 without it.
- Frogger: 10.3 crossings, level with the rule bot's 10.25, against 1.0 for the untrained base model.
- Pac-Man: 57.0 of 98 pellets, about 60% of the rule bot's 94.1.
The odd result is size. Jeff-2B scores higher on benchmarks but plays worse (6.0 Frogger crossings, -0.9 Doom kills), which the author attributes to the untrained 2B already being more risk-averse and says training made it worse. A 2B model too cautious to shoot demons is a very relatable failure mode.
What you can reuse
- Fine-tune past the zero-shot ceiling. The author uses the 0.8B for voice navigation in one of their apps. A fine-tune on about 11,000 app-specific examples took about half an hour on one GPU and moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision on an M4 Max.
- Reason in code, decide with the model. Jeff is a classifier, not a planner. Asked to forecast ("a car arrives in 2 turns"), it does no better than random. Precompute consequences in code and let the model choose between them.
- Wording is a hyperparameter. Giving Frogger's goal option the same phrasing as every other forward option took one episode from 15 crossings to 23. Jev's own Doom prompt, a raw bearing number plus an aiming rule, does not work for any Jeff model.
- The recipe is public.
autojev-mix, the training script, andautojev-evaluateare in the repo, so you can train your own student on your own labels.
Key Takeaways
- Jeff is a drop-in local alternative to Jev's request format: three Apache 2.0 models (0.8B, 2B, Gemma 4 E2B) that return calibrated option probabilities in one forward pass.
- The overall score ties Jev, the profile does not: 96.3% vs 77.0% on Financial PhraseBank, but 68.0% vs 94.3% on BBH. Use it for classification, routing, and grounding checks, not reasoning.
- The comparison is cross-study: Jev and AutoJev numbers are published figures on a different sample, so treat 83.1% vs 83.0% as "same band," not "beats."
- Local latency is the real advantage: 22 ms on an RTX PRO 6000, 28 ms on an M4 Max, versus 114 to 212 ms for Jev over its API in published Doom runs.
- The whole thing was built at home: one workstation GPU for about 2 hours of training, open-model synthetic data from two DGX Sparks, no cloud GPUs and no closed-model output in training.
Sources: firelex/jeff on GitHub, Jeff-Qwen3.5-0.8B model card, Jeff-Qwen3.5-2B model card, Show HN post and author comments, AutoJev on GitHub, TypeSafe, jev-plays-doom