PostHog · 2026-09-29 · notable
Jeeves — PostHog's 9B decision model reasons before it picks
Jeeves is PostHog's open 9B Jev-compatible decision model. It thinks before it answers and scores 0.935 on JevBench's public tiers against Jev's 0.866. Weights, training code and data are MIT-licensed.
An open decision model that writes a reasoning chain first, then returns calibrated probabilities for each option.
Key specs
| Jev bench public | 0.935 |
|---|---|
| Jev bench hard | 0.865 |
What is it?
Jeeves is a 9B decision classifier from PostHog that takes Jev's request format — yes/no, multiple-choice and rating questions — and returns a probability for each option. Unlike most Jev-like models it reasons before it decides. The weights are on Hugging Face and the repo ships the full training code and data.
How does it work?
The model is Qwen3.5-9B with a LoRA adapter and a pointer head that scores each option at a special decide token after the reasoning chain. Training used supervised fine-tuning on 19,126 questions, then CISPO reinforcement learning, then a fitted temperature for calibration. A diffusion drafter adapted to Qwen3.5's Gated DeltaNet layers speeds up the reasoning chain by about 1.6×.
Why does it matter?
Pipelines often use a Jev-like model for fast, calibrated answers and fall back to a large reasoning model when accuracy matters. Jeeves tries to do both in one self-hosted model: 0.889 on its out-of-domain test split against 0.857 for Jev. The trade-offs are stated plainly — it trails Jev on MMLU (0.793 vs 0.900), and full thinking takes a 3.3 s median on one H100, so answers can be capped or skip thinking.
Who is it for?
ML engineers building classification and routing pipelines
Try it
hf download PostHog/jeeves --local-dir jeeves-weights