Overview
Jeff is a family of small zero-shot decision models you run next to your own code. You describe a situation and list the options in plain words — support queues, user intents, moderation labels, voice commands, game moves — and Jeff returns a calibrated probability for each option from a single forward pass. There is no generated text and nothing to parse, and the options do not need to appear in the training data: you describe them and Jeff picks.
Three checkpoints are published on Hugging Face under Apache 2.0: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B. The median time per decision is 22 ms for the 0.8B on an RTX PRO 6000 and 28 ms on an Apple M4 Max through MLX. Across 4,599 questions from five public benchmarks the 2B scores 83.1% overall against Jev's published 83.0%, with most of that coming from classification and grounding tasks — on reasoning-heavy BBH and JudgeBench it stays well below the larger models, as the authors expect at this size.

Jeff accepts the same request format as TypeSafe's Jev but is an independent project, not affiliated with or endorsed by TypeSafe. It began as a fork of Denis Yarats's MIT-licensed AutoJev recipe and adds small student models, a local synthetic-data pipeline with a leak filter, MLX serving on Apple silicon and a set of game tests. Everything was trained on one RTX PRO 6000 workstation GPU — about 2 hours for the 0.8B and 3.5 hours for the 2B — with all synthetic data written by the open Qwen3.8-Flash-Next on two DGX Sparks. Within structured output it sits beside Laya and SemIf, which also replace generation with a single scoring pass.
What it does
- Three question types: `choice` picks one of up to 255 options, `noul` returns a yes/no probability, and `score` places a value on a scale you describe
- Several independent questions in one request are answered together
- Calibrated probabilities from one fitted temperature, plus the chosen option and a confidence per answer
- An HTTP server (`jeff-serve`) with a Jev-style `/v1/systemone` endpoint, on PyTorch for NVIDIA or CPU and on MLX for Apple silicon
- AutoJev-based training scripts for your own domain fine-tunes, with a leak filter and a training dashboard
- A game harness for Doom, Frogger and Pac-Man to test zero-shot choices outside the benchmarks
Getting started
Jeff uses uv for its environment and pulls weights from Hugging Face. The Apple-silicon MLX backend runs the Qwen checkpoints only.
Install and download a checkpoint
uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8bStart the server
Use the PyTorch backend on an NVIDIA GPU or CPU, or the MLX backend on a Mac, which is much faster there.
# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
# Apple silicon (MLX, Qwen models only)
uv sync --extra mac
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serveAsk a question
Describe the state and the options in words; each answer comes back with a probability per option, the chosen option and a confidence.
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"model": "jeff-latest",
"state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
"questions": {
"route": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Routing support tickets or requests to the right queue locally, without an API call per decision
- Moderation labels, intent detection and voice-command classification where tens of milliseconds matter

- A fast decision step inside an agent loop, with the reasoning kept in code and the choice left to Jeff
- A base to fine-tune for one narrow task — the authors' voice-navigation run went from 31.7% to 95.8% held-out accuracy in about half an hour on one GPU
How Jeff compares
Jeff alongside other open-source structured output tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Laya | ★ 27.7k | An Apache-2.0 decision model that answers choice, score and boolean questions about text in one forward pass, returning calibrated probabilities instead of generated JSON. |
| Guidance | ★ 21.8k | A programming model that interleaves generation, prompting, and control logic to constrain output and enforce formats like JSON or regex patterns. |
| Outlines | ★ 15.9k | A library for structured generation that constrains an LLM's token output to match a JSON schema, regex, or grammar so the result is always valid. |
| Instructor | ★ 14k | A library that wraps an LLM client to return data validated against a schema, retrying automatically on invalid output, with SDKs in several languages. |
| BAML | ★ 9.3k | A domain-specific language for defining LLM functions with typed schemas, parsing flexible model output into reliable structured data across many languages. |
| Marvin | ★ 6.2k | A Python toolkit from Prefect for turning LLM calls into typed functions that extract, classify, and cast text into structured Python objects. |
| SemIf | ★ 4.5k | A decision baseline that reads typed option probabilities directly from an open model in one forward pass, instead of generating and parsing JSON. |
| Jeff | — | Tiny local decision models that speak Jev's request format and answer in about 22 milliseconds |