Overview
Jeeves is a decision classifier that answers questions in Jev's request format — yes/no (`noul`), multiple choice (`choice`) and rating (`score`) — and returns a calibrated probability for each option. Where most Jev-like models answer straight from the prompt, Jeeves first writes a reasoning chain and only then decides. The authors built it because Jev-like models give calibrated probabilities but at low accuracy, so many pipelines keep a full reasoning model as a fallback.
Under the hood it is Qwen3.5-9B with a LoRA adapter and a pointer head. After the reasoning chain the questions are repeated, and the pointer head scores each option by comparing the hidden state at a `<decide>` token with the hidden state at the end of each option; a softmax with a temperature fitted on the dev set turns those scores into probabilities. Training ran supervised fine-tuning on 19,126 questions from 12 public datasets plus synthetic policy data, then CISPO reinforcement learning, then calibration. A diffusion drafter, adapted from Orthrus to work with Qwen3.5's Gated DeltaNet layers, speeds up the reasoning chain by about 1.6× at block size 4.
PostHog publishes the 9B weights on Hugging Face and the full training code and train/dev/test data under the MIT license. The README reports 0.889 on its out-of-domain test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's 231 public items against Jev's 0.866. It trails Jev on knowledge questions (MMLU 0.793 against 0.900), and full thinking is slow at the tail — a 3.3 s median and 17.1 s p90 on one H100 — so the server lets you cap or skip thinking per request. Within structured output it sits beside Jeff, Laya and SemIf, which answer the same kind of question without a reasoning step.
What it does
- Answers `noul`, `choice` and `score` questions in the same request through a Jev-compatible `/v1/systemone` endpoint
- Reasons before deciding, with `think`, `max_think` and `nothink_threshold` options to trade accuracy for latency
- About 0.3 s per request without thinking and a 3.3 s median with it on one H100
- Block-4 diffusion drafter for faster reasoning chains, about 960 chain tokens per second across eight batched questions
- `jeeves_sdk`, a drop-in replacement for Jev's `typesafe-sdk` Python client that can also return each question's reasoning text
- Full reproduction pipeline: dataset prep, SFT, CISPO, calibration, LoRA fusing and drafter training
Getting started
Jeeves needs Python 3.12 and a CUDA GPU (Hopper for the FP8 kernel). The released weights come from Hugging Face.
Install and download the weights
pip install -r requirements.txt
hf download PostHog/jeeves --local-dir jeeves-weightsStart the server with the drafter
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009Ask a question in Jev's format
Each answer comes back with a probability per option; `max_think` caps how long each reasoning chain can run.
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice. Please help.",
"questions": {
"billing": {"type": "noul", "instructions": "Is this about billing?"}
},
"options": {"max_think": 512}}'Or use the Python SDK
The SDK connects to http://127.0.0.1:8009 by default (or `JEEVES_BASE_URL`) and needs no API key.
pip install ./sdkCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Replacing a reasoning-model fallback in a classification pipeline with one self-hosted decision model
- Routing and triaging support messages where a calibrated probability per team or label is needed
- Policy and rule checks where the model must weigh several conditions before answering yes or no
- A research base for training reasoning decision models, with the full recipe and data published
How Jeeves compares
Jeeves alongside other open-source structured output tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Laya | ★ 28.4k | An Apache-2.0 decision model that answers choice, score and boolean questions about text in one forward pass, returning calibrated probabilities instead of generated JSON. |
| Guidance | ★ 21.8k | A programming model that interleaves generation, prompting, and control logic to constrain output and enforce formats like JSON or regex patterns. |
| Outlines | ★ 15.9k | A library for structured generation that constrains an LLM's token output to match a JSON schema, regex, or grammar so the result is always valid. |
| Instructor | ★ 14k | A library that wraps an LLM client to return data validated against a schema, retrying automatically on invalid output, with SDKs in several languages. |
| BAML | ★ 9.4k | A domain-specific language for defining LLM functions with typed schemas, parsing flexible model output into reliable structured data across many languages. |
| Marvin | ★ 6.2k | A Python toolkit from Prefect for turning LLM calls into typed functions that extract, classify, and cast text into structured Python objects. |
| SemIf | ★ 4.6k | A decision baseline that reads typed option probabilities directly from an open model in one forward pass, instead of generating and parsing JSON. |
| Jeeves | — | A 9B Jev-compatible decision model that thinks before it picks, with open weights, training code and data |