█

AI/TLDR

Jeeves

An open 9B decision model from PostHog that writes a reasoning chain before it answers typed questions with calibrated probabilities, behind a Jev-compatible API, released September 2026.

Jeeves (open decision models)Open weightsGenerally available
Released
29 Sep 2026
Parameters
9B (Qwen3.5-9B base with the LoRA merged, plus a 256-dim pointer head)
License
Apache-2.0 (weights), MIT (code)

Overview

Jeeves is a 9B open decision model published by PostHog, with the code on GitHub and the weights on Hugging Face on September 29, 2026. Like TypeSafe's Jev, a request carries a state (text or JSON) plus typed questions — yes/no (noul), multiple-choice (choice) and rating (score) — and every answer comes back as a typed value with a probability for each option. What sets Jeeves apart is that it writes a reasoning chain for each question before it decides.

The model is Qwen3.5-9B with a LoRA adapter (r=16 on all projections) and a pointer head. After the reasoning chain closes, the question and its options are repeated and a special decide token is appended; the head scores each option with a scaled dot product between the hidden state at that token and the hidden state at the option's closing tag, and a temperature fitted on the dev set turns the scores into calibrated probabilities. Training ran supervised fine-tuning on 19,126 questions from 12 public datasets and synthetic policy data, then CISPO reinforcement learning on 9,992 questions (the released checkpoint is step 402 of a 624-step schedule), then calibration.

In the README's results, Jeeves scores 0.889 on its out-of-domain and held-out test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's 231 public items against 0.866 for Jev (0.865 against 0.730 on the hard tier). The Kev and Jev figures are the ones Kev publishes, and outside JevBench they use different items from the same sources. Jeeves trails Jev on knowledge questions (MMLU 0.793 against 0.900, MMLU-Pro 0.739 against 0.840).

Thinking costs time: on one H100 the README reports a 3.3 s median and 17.1 s p90 with full chains, 2.0 s median with chains capped at 768 tokens and no-think for confident answers, and about 0.3 s with thinking off (0.804 on the test split against 0.840 with thinking). A diffusion drafter adapted to Qwen3.5's Gated DeltaNet layers speeds up chain decoding 1.6× at block 4. The repo ships the server (CUDA, with an FP8 kernel for Hopper), a drop-in replacement for Jev's Python SDK, and the full training code and data pipeline.

Released2026-09-29
LicenseApache-2.0 (weights), MIT (code)
WeightsOpen weights
Parameters9B (Qwen3.5-9B base with the LoRA merged, plus a 256-dim pointer head)
ArchitectureQwen3.5-9B (Gated DeltaNet + gated attention) with a LoRA (r=16) merged in and a pointer head that scores each option at a decide token after the reasoning chain; block-4 and block-8 diffusion drafters for speculative decoding
ModalitiesText
StatusGenerally available

Benchmarks

Jeeves against Kev-9B and TypeSafe's Jev, as published in the Jeeves README (accuracy with thinking, greedy, 2,560-token cap). Kev-9B and Jev figures are the ones Kev publishes; the Kev-9B JevBench cells are Kev-8B (Qwen3), since no Kev-9B JevBench result is published.

BenchmarkKev-9BJevJeeves
Test overall (out-of-domain and held-out)0.8220.8570.889
Transfer overall (MMLU-Pro and buried state)0.5790.80.746
JevBench overall (231 public items)0.7150.8660.935
QNLI0.9250.9250.913
SciQ0.9630.9880.991
TweetEval offensive0.7750.8130.813
PAWS0.7630.7880.875
MMLU0.7380.90.793
Emotion0.60.5880.647
Held-out rule structures0.8960.8851
Contrastive policies0.90.9631
MMLU-Pro (10-way)0.5150.840.739
Buried state0.740.70.759
Unknowable answered at p ≥ 0.9 (lower is better)00.090.055
JevBench hard (111 public items)0.4510.730.865
JevBench ECE, public items (lower is better)—0.0490.037

Comparison source ↗

This model's scores

  1. JevBench, public tiers (231 items)93.5%
  2. Test overall (out-of-domain and held-out)88.9%
  3. JevBench hard (111 public items)86.5%
  4. MMLU-Pro (10-way)73.9%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Strengths

  • 0.935 on JevBench's 231 public items against 0.866 for Jev, and 0.865 against 0.730 on the hard tier
  • 0.889 on out-of-domain and held-out test data it was never trained on, against 0.857 for Jev and 0.822 for Kev-9B
  • Reasoning is tunable per request: cap chains with max_think, skip thinking when the no-think answer is confident, or turn it off for about 0.3 s answers
  • Jev-compatible /v1/systemone server and a drop-in replacement for Jev's Python SDK
  • Full training code, dataset build scripts and pinned dataset revisions; a from-scratch reproduction matched the released checkpoint within noise

Best for

  • Reach for it when a pipeline uses a fast Jev-like classifier but falls back to a large reasoning model for hard cases, and you want one self-hosted model that does both.
  • Reach for it to route, triage or score tickets, emails and policy checks where your code should act only on confident, calibrated answers.
  • Reach for it when you want to study or retrain a reasoning decision model end to end, since the SFT, CISPO and drafter training code are all in the repo.
  • Look elsewhere for knowledge-heavy questions, where it trails Jev, or for latency-critical paths that cannot afford a multi-second reasoning tail.

How to access

ProviderModel ID
Hugging Face (weights) ↗PostHog/jeeves

FAQ

What makes Jeeves different from other Jev-like decision models?

It reasons first. For each question Jeeves writes a reasoning chain, then a pointer head scores every option and a fitted temperature turns the scores into calibrated probabilities. Other Jev-like models such as Kev answer directly from the prompt.

How does Jeeves compare with TypeSafe's Jev?

In the Jeeves README it scores 0.935 on JevBench's 231 public items against 0.866 for Jev, and 0.889 on its out-of-domain and held-out test split against 0.857. Jev leads on knowledge questions (MMLU 0.900 against 0.793, MMLU-Pro 0.840 against 0.739). The Jev figures are the ones Kev publishes, and outside JevBench they use different items from the same sources.

How fast is Jeeves?

On one H100 over 325 dev questions, full thinking has a 3.3 s median and 17.1 s p90 latency; capping chains at 768 tokens and skipping thinking for confident answers gives a 2.0 s median; with thinking off it answers in about 0.3 s, at lower accuracy.

What license is Jeeves released under?

The Jeeves code is MIT. The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license, according to the Hugging Face model card.

What hardware does it need?

Serving needs a CUDA GPU. The FP8 kernel needs a Hopper GPU; on other cards the server runs in bf16 with --no-fp8.