█

AI/TLDR

Jeeves

A 9B Jev-compatible decision model that thinks before it picks, with open weights, training code and data

Structured OutputOpen source
Updated
29 Sep 2026
Language
Python
License
MIT
Coverage
1 story
$pip install -r requirements.txt

What's new

29 Sep 2026

PostHog released Jeeves: a 9B Jev-compatible decision model trained with SFT and CISPO to reason before it decides, with open weights, training code and data. It scores 0.935 on JevBench's public tiers against Jev's 0.866.

Latest news

Overview

Jeeves is a decision classifier that answers questions in Jev's request format — yes/no (`noul`), multiple choice (`choice`) and rating (`score`) — and returns a calibrated probability for each option. Where most Jev-like models answer straight from the prompt, Jeeves first writes a reasoning chain and only then decides. The authors built it because Jev-like models give calibrated probabilities but at low accuracy, so many pipelines keep a full reasoning model as a fallback.

Under the hood it is Qwen3.5-9B with a LoRA adapter and a pointer head. After the reasoning chain the questions are repeated, and the pointer head scores each option by comparing the hidden state at a `<decide>` token with the hidden state at the end of each option; a softmax with a temperature fitted on the dev set turns those scores into probabilities. Training ran supervised fine-tuning on 19,126 questions from 12 public datasets plus synthetic policy data, then CISPO reinforcement learning, then calibration. A diffusion drafter, adapted from Orthrus to work with Qwen3.5's Gated DeltaNet layers, speeds up the reasoning chain by about 1.6× at block size 4.

PostHog publishes the 9B weights on Hugging Face and the full training code and train/dev/test data under the MIT license. The README reports 0.889 on its out-of-domain test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's 231 public items against Jev's 0.866. It trails Jev on knowledge questions (MMLU 0.793 against 0.900), and full thinking is slow at the tail — a 3.3 s median and 17.1 s p90 on one H100 — so the server lets you cap or skip thinking per request. Within structured output it sits beside Jeff, Laya and SemIf, which answer the same kind of question without a reasoning step.

What it does

  • Answers `noul`, `choice` and `score` questions in the same request through a Jev-compatible `/v1/systemone` endpoint
  • Reasons before deciding, with `think`, `max_think` and `nothink_threshold` options to trade accuracy for latency
  • About 0.3 s per request without thinking and a 3.3 s median with it on one H100
  • Block-4 diffusion drafter for faster reasoning chains, about 960 chain tokens per second across eight batched questions
  • `jeeves_sdk`, a drop-in replacement for Jev's `typesafe-sdk` Python client that can also return each question's reasoning text
  • Full reproduction pipeline: dataset prep, SFT, CISPO, calibration, LoRA fusing and drafter training

Getting started

Jeeves needs Python 3.12 and a CUDA GPU (Hopper for the FP8 kernel). The released weights come from Hugging Face.

Install and download the weights

bashbash
pip install -r requirements.txt
hf download PostHog/jeeves --local-dir jeeves-weights

Start the server with the drafter

bashbash
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009

Ask a question in Jev's format

Each answer comes back with a probability per option; `max_think` caps how long each reasoning chain can run.

bashbash
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
  "state": "I was charged twice. Please help.",
  "questions": {
    "billing": {"type": "noul", "instructions": "Is this about billing?"}
  },
  "options": {"max_think": 512}}'

Or use the Python SDK

The SDK connects to http://127.0.0.1:8009 by default (or `JEEVES_BASE_URL`) and needs no API key.

bashbash
pip install ./sdk

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Replacing a reasoning-model fallback in a classification pipeline with one self-hosted decision model
  • Routing and triaging support messages where a calibrated probability per team or label is needed
  • Policy and rule checks where the model must weigh several conditions before answering yes or no
  • A research base for training reasoning decision models, with the full recipe and data published

How Jeeves compares

Jeeves alongside other open-source structured output tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Laya★ 28.4kAn Apache-2.0 decision model that answers choice, score and boolean questions about text in one forward pass, returning calibrated probabilities instead of generated JSON.
Guidance★ 21.8kA programming model that interleaves generation, prompting, and control logic to constrain output and enforce formats like JSON or regex patterns.
Outlines★ 15.9kA library for structured generation that constrains an LLM's token output to match a JSON schema, regex, or grammar so the result is always valid.
Instructor★ 14kA library that wraps an LLM client to return data validated against a schema, retrying automatically on invalid output, with SDKs in several languages.
BAML★ 9.4kA domain-specific language for defining LLM functions with typed schemas, parsing flexible model output into reliable structured data across many languages.
Marvin★ 6.2kA Python toolkit from Prefect for turning LLM calls into typed functions that extract, classify, and cast text into structured Python objects.
SemIf★ 4.6kA decision baseline that reads typed option probabilities directly from an open model in one forward pass, instead of generating and parsing JSON.
Jeeves—A 9B Jev-compatible decision model that thinks before it picks, with open weights, training code and data