Overview
Jeff is a set of three small open decision models published by Firelex on GitHub and Hugging Face on September 28, 2026: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B. The weights are Apache 2.0 and the code is MIT. You describe a situation and list the options in plain words, and Jeff returns a calibrated probability for each option from a single forward pass, with no generated text to parse. It accepts the same request format as TypeSafe's Jev (`POST /v1/systemone` with `choice`, `noul` and `score` questions) but is an independent project, not affiliated with or endorsed by TypeSafe.
Jeff began as a fork of Denis Yarats's MIT-licensed AutoJev recipe and keeps its core design: one forward pass per decision, a trained answer readout and a fitted temperature. The models are full-weight fine-tunes trained for one epoch on a single RTX PRO 6000 workstation GPU (about 2 hours for the 0.8B and 3.5 hours for the 2B). The synthetic part of the training data was written by an open model, Qwen3.8-Flash-Next, on two DGX Sparks, and every training question is checked against the benchmark panel by a leak filter.
On 4,599 questions from five public benchmarks, Jeff-Qwen3.5-2B scores 82.0%, Jeff-Gemma4-E2B 81.6% and Jeff-Qwen3.5-0.8B 79.1%, against 83.0% published for Jev, which was measured on a different sample of the same benchmarks. Jeff's score comes from classification and grounding: all three models beat Jev's published figures on Financial PhraseBank and RAGTruth, while they stay well below it on the reasoning-heavy BBH, JudgeBench and JevBench hard tier. The authors state plainly that models this small do not reason.
Version 1.1 of the two Qwen models followed on September 29, 2026. It raises the option limit from 26 to 254 and lifts the 0.8B from 40.3% to 94.7% on a long-list test, and it lowers calibration error to 0.021 (0.8B) and 0.026 (2B). The 2B's benchmark score moved from 83.1% to 82.0%, mostly on JudgeBench. Jeff-Gemma4-E2B was not retrained, stays at v1.0 and picks reliably among at most 26 options.
As a zero-shot test, the authors had the models play Doom, Frogger and Pac-Man from options that describe each move's consequence in words. With v1.0 over 20 episodes, Jeff-Qwen3.5-0.8B matched the hand-coded rule bot in Doom (6.55 kills) and Frogger (10.3 crossings against 10.25) and ate 57.0 of 98 Pac-Man pellets. The 2B plays worse than the 0.8B despite scoring higher on the benchmarks, which the README flags as an open question.
When zero-shot is not enough, the README shows fine-tuning. A voice-navigation fine-tune on about 11,000 app-specific examples moved held-out accuracy from 31.7% to 95.8% in about half an hour on one GPU. Jeff-Qwen3.5-0.8B-Chess, trained on 600,000 Stockfish-labelled Lichess positions in about 3.5 hours, solves 55.8% of 1,000 held-out puzzles against 15.5% zero-shot. The authors put its strength at about 1,000 Elo with no search.

The repository includes a server for NVIDIA GPUs and CPU (PyTorch) and for Apple silicon through MLX (Qwen models only). Median time per decision is 22 ms for the 0.8B on an RTX PRO 6000, 28 ms on an Apple M4 Max and 463 ms on a 32-thread CPU. The models are English and text only.
| Released | 2026-09-28 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | Three models: Jeff-Qwen3.5-0.8B (0.8B), Jeff-Qwen3.5-2B (2B) and Jeff-Gemma4-E2B (2B effective, 4.6B stored) |
| Architecture | Full-weight fine-tunes of Qwen3.5-0.8B, Qwen3.5-2B and Gemma 4 E2B that read out a probability for each option letter in a single forward pass, with one fitted temperature for calibration; training starts from the open AutoJev recipe |
| Modalities | Text |
Benchmarks

Jeff against the published Jev and AutoJev-27B figures, as published in the Jeff README (accuracy, %). The Qwen columns are v1.1 and Jeff-Gemma4-E2B is v1.0; the Jev and AutoJev figures were measured on a different sample of the same benchmarks. JevBench hard tier (105 items) is scored separately from the five-benchmark overall.
| Benchmark | Jeff-Qwen3.5-0.8B | Jeff-Qwen3.5-2B | Jeff-Gemma4-E2B | Jev (published) | AutoJev-27B (published) |
|---|---|---|---|---|---|
| Overall (5 benchmarks) | 79.1% | 82% | 81.6% | 83% | 84.9% |
| BBH | 64.9% | 68.7% | 66.4% | 94.3% | 82.8% |
| Financial PhraseBank | 95.7% | 94.7% | 96.1% | 77% | 84.2% |
| JudgeBench | 63.1% | 59.4% | 60.6% | 78.6% | 78.9% |
| RAGTruth | 85.6% | 87.7% | 87.4% | 77.3% | 88.9% |
| WinoGrande | 69% | 78.8% | 77.4% | 90.7% | 83.3% |
| JevBench hard tier (separate) | 46.7% | 57.1% | 48.6% | 73.3% | 70.3% |
This model's scores
- Overall accuracy, 5 public benchmarks (Jeff-Qwen3.5-2B, v1.1)82%
- Overall accuracy, 5 public benchmarks (Jeff-Gemma4-E2B)81.6%
- Overall accuracy, 5 public benchmarks (Jeff-Qwen3.5-0.8B, v1.1)79.1%
- Financial PhraseBank accuracy (Jeff-Qwen3.5-0.8B, v1.1)95.7%
- RAGTruth accuracy (Jeff-Qwen3.5-2B, v1.1)87.7%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- A calibrated probability for every option from a single forward pass: 22 ms per decision for Jeff-Qwen3.5-0.8B on an RTX PRO 6000 and 28 ms on an Apple M4 Max
- Jeff-Qwen3.5-2B scores 82.0% across five public benchmarks against Jev's published 83.0%, and all three models beat Jev's published figures on Financial PhraseBank and RAGTruth
- Same request format as Jev, served locally from one command on NVIDIA GPUs, CPU or Apple silicon
- v1.1 Qwen models choose among up to 254 options, with calibration error of 0.021 (0.8B) and 0.026 (2B)
- Apache-2.0 weights and MIT code, trained entirely on local hardware with a documented fine-tuning path (voice navigation: 31.7% to 95.8% held-out accuracy)
Best for
- Reach for it when code needs a fast, local pick between options you describe in words: support queues, user intents, moderation labels, voice commands.
- Reach for it when you already send Jev-style requests and want a small self-hosted model that answers in tens of milliseconds.
- Reach for it as a base for a short domain fine-tune when zero-shot accuracy is not enough.
- Look elsewhere for multi-step reasoning or forecasting, where the README reports Jeff well below Jev on BBH, JudgeBench and JevBench and no better than random when asked to predict what happens next.
How to access
| Provider | Model ID |
|---|---|
| Hugging Face (weights, 0.8B) ↗ | mstrasser/Jeff-Qwen3.5-0.8B |
| Hugging Face (weights, 2B) ↗ | mstrasser/Jeff-Qwen3.5-2B |
| Hugging Face (weights, Gemma 4 E2B) ↗ | mstrasser/Jeff-Gemma4-E2B |
FAQ
What does Jeff do differently from a chat model?
Jeff never writes text. You describe a situation and list the options in words, and it returns a calibrated probability for each option from one forward pass, plus the chosen option and a confidence. Question types are choice, noul (yes/no) and score.
How does Jeff compare with TypeSafe's Jev?
In the Jeff README, Jeff-Qwen3.5-2B scores 82.0% across five public benchmarks against 83.0% published for Jev. All three Jeff models beat Jev's published figures on Financial PhraseBank and RAGTruth and stay well below it on BBH, JudgeBench, WinoGrande and the JevBench hard tier. Jev's figures were measured on a different sample of the same benchmarks, and Jeff is not affiliated with TypeSafe.
Which Jeff model should I use?
The README calls the 0.8B the sweet spot for fast option picking: 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max. The 2B scores higher on the benchmarks but is more cautious and plays the test games worse. Jeff-Gemma4-E2B has no MLX backend on a Mac and, at v1.0, picks reliably among at most 26 options.
What changed in Jeff v1.1?
Released September 29, 2026 for the two Qwen models, v1.1 raises the option limit from 26 to 254, moves the 0.8B from 40.3% to 94.7% on a long-list test and improves calibration error to 0.021 (0.8B) and 0.026 (2B). The 2B's benchmark score moved from 83.1% to 82.0%. v1.0 stays on Hugging Face as revision v1.0.
Is Jeff open source?
Yes. The weights are Apache 2.0 and the code is MIT, including the AutoJev code it builds on. The training data is not released; its sources and licences are listed in the repository.


