Overview
Jeeves is a 9B open decision model published by PostHog, with the code on GitHub and the weights on Hugging Face on September 29, 2026. Like TypeSafe's Jev, a request carries a state (text or JSON) plus typed questions — yes/no (noul), multiple-choice (choice) and rating (score) — and every answer comes back as a typed value with a probability for each option. What sets Jeeves apart is that it writes a reasoning chain for each question before it decides.
The model is Qwen3.5-9B with a LoRA adapter (r=16 on all projections) and a pointer head. After the reasoning chain closes, the question and its options are repeated and a special decide token is appended; the head scores each option with a scaled dot product between the hidden state at that token and the hidden state at the option's closing tag, and a temperature fitted on the dev set turns the scores into calibrated probabilities. Training ran supervised fine-tuning on 19,126 questions from 12 public datasets and synthetic policy data, then CISPO reinforcement learning on 9,992 questions (the released checkpoint is step 402 of a 624-step schedule), then calibration.
In the README's results, Jeeves scores 0.889 on its out-of-domain and held-out test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's 231 public items against 0.866 for Jev (0.865 against 0.730 on the hard tier). The Kev and Jev figures are the ones Kev publishes, and outside JevBench they use different items from the same sources. Jeeves trails Jev on knowledge questions (MMLU 0.793 against 0.900, MMLU-Pro 0.739 against 0.840).
Thinking costs time: on one H100 the README reports a 3.3 s median and 17.1 s p90 with full chains, 2.0 s median with chains capped at 768 tokens and no-think for confident answers, and about 0.3 s with thinking off (0.804 on the test split against 0.840 with thinking). A diffusion drafter adapted to Qwen3.5's Gated DeltaNet layers speeds up chain decoding 1.6× at block 4. The repo ships the server (CUDA, with an FP8 kernel for Hopper), a drop-in replacement for Jev's Python SDK, and the full training code and data pipeline.
| Released | 2026-09-29 |
|---|---|
| License | Apache-2.0 (weights), MIT (code) |
| Weights | Open weights |
| Parameters | 9B (Qwen3.5-9B base with the LoRA merged, plus a 256-dim pointer head) |
| Architecture | Qwen3.5-9B (Gated DeltaNet + gated attention) with a LoRA (r=16) merged in and a pointer head that scores each option at a decide token after the reasoning chain; block-4 and block-8 diffusion drafters for speculative decoding |
| Modalities | Text |
| Status | Generally available |
Benchmarks
Jeeves against Kev-9B and TypeSafe's Jev, as published in the Jeeves README (accuracy with thinking, greedy, 2,560-token cap). Kev-9B and Jev figures are the ones Kev publishes; the Kev-9B JevBench cells are Kev-8B (Qwen3), since no Kev-9B JevBench result is published.
| Benchmark | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out) | 0.822 | 0.857 | 0.889 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.8 | 0.746 |
| JevBench overall (231 public items) | 0.715 | 0.866 | 0.935 |
| QNLI | 0.925 | 0.925 | 0.913 |
| SciQ | 0.963 | 0.988 | 0.991 |
| TweetEval offensive | 0.775 | 0.813 | 0.813 |
| PAWS | 0.763 | 0.788 | 0.875 |
| MMLU | 0.738 | 0.9 | 0.793 |
| Emotion | 0.6 | 0.588 | 0.647 |
| Held-out rule structures | 0.896 | 0.885 | 1 |
| Contrastive policies | 0.9 | 0.963 | 1 |
| MMLU-Pro (10-way) | 0.515 | 0.84 | 0.739 |
| Buried state | 0.74 | 0.7 | 0.759 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0 | 0.09 | 0.055 |
| JevBench hard (111 public items) | 0.451 | 0.73 | 0.865 |
| JevBench ECE, public items (lower is better) | — | 0.049 | 0.037 |
This model's scores
- JevBench, public tiers (231 items)93.5%
- Test overall (out-of-domain and held-out)88.9%
- JevBench hard (111 public items)86.5%
- MMLU-Pro (10-way)73.9%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- 0.935 on JevBench's 231 public items against 0.866 for Jev, and 0.865 against 0.730 on the hard tier
- 0.889 on out-of-domain and held-out test data it was never trained on, against 0.857 for Jev and 0.822 for Kev-9B
- Reasoning is tunable per request: cap chains with max_think, skip thinking when the no-think answer is confident, or turn it off for about 0.3 s answers
- Jev-compatible /v1/systemone server and a drop-in replacement for Jev's Python SDK
- Full training code, dataset build scripts and pinned dataset revisions; a from-scratch reproduction matched the released checkpoint within noise
Best for
- Reach for it when a pipeline uses a fast Jev-like classifier but falls back to a large reasoning model for hard cases, and you want one self-hosted model that does both.
- Reach for it to route, triage or score tickets, emails and policy checks where your code should act only on confident, calibrated answers.
- Reach for it when you want to study or retrain a reasoning decision model end to end, since the SFT, CISPO and drafter training code are all in the repo.
- Look elsewhere for knowledge-heavy questions, where it trails Jev, or for latency-critical paths that cannot afford a multi-second reasoning tail.
How to access
| Provider | Model ID |
|---|---|
| Hugging Face (weights) ↗ | PostHog/jeeves |
FAQ
What makes Jeeves different from other Jev-like decision models?
It reasons first. For each question Jeeves writes a reasoning chain, then a pointer head scores every option and a fitted temperature turns the scores into calibrated probabilities. Other Jev-like models such as Kev answer directly from the prompt.
How does Jeeves compare with TypeSafe's Jev?
In the Jeeves README it scores 0.935 on JevBench's 231 public items against 0.866 for Jev, and 0.889 on its out-of-domain and held-out test split against 0.857. Jev leads on knowledge questions (MMLU 0.900 against 0.793, MMLU-Pro 0.840 against 0.739). The Jev figures are the ones Kev publishes, and outside JevBench they use different items from the same sources.
How fast is Jeeves?
On one H100 over 325 dev questions, full thinking has a 3.3 s median and 17.1 s p90 latency; capping chains at 768 tokens and skipping thinking for confident answers gives a 2.0 s median; with thinking off it answers in about 0.3 s, at lower accuracy.
What license is Jeeves released under?
The Jeeves code is MIT. The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license, according to the Hugging Face model card.
What hardware does it need?
Serving needs a CUDA GPU. The FP8 kernel needs a Hopper GPU; on other cards the server runs in bf16 with --no-fp8.