github.com via Hacker News

PostHog Ships Jeeves, 9B Decision Model Beating Jev on JevBench

Open Source ai-business

TL;DR

  • Jeeves scores 0.935 on the 231-item JevBench public tier and 0.889 on out-of-domain test data, beating Jev's 0.866 and 0.857.
  • The model is a LoRA-rank-16 adaptation of Qwen3.5-9B with a custom pointer head, trained with SFT then CISPO reinforcement learning.
  • Knowledge tasks regress against Jev: MMLU drops to 0.793 from 0.900 and MMLU-Pro to 0.739 from 0.840, per the limitations section.

Jeeves, a 9B decision classifier released on PostHog's GitHub, scores 0.935 on the 231-item JevBench public tier, up from Jev's 0.866 and Kev-9B's 0.715, and 0.889 on out-of-domain test data it was never trained on.

The recipe is a LoRA-rank-16 adaptation of Qwen3.5-9B with a custom pointer head that scores options via scaled dot product between query projections at the decide token and key projections at option end tokens. Training runs in three stages: supervised fine-tuning across 596 steps on 8 GPUs over 19,126 questions from 12 public datasets, then CISPO reinforcement learning stopped at step 402 with 8 rollouts per question capped at 2,560 thinking tokens, then a single temperature parameter fitted on a development set. The repository describes it as "a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides."

Latency splits sharply on whether reasoning runs. Without thinking, the model returns in about 0.3 seconds; with full thinking chains, median climbs to 3.3s and p90 stretches to 17.1s over roughly 1,138 reasoning tokens. A block-4 diffusion drafter adapted from the Orthrus architecture, modified to support Gated DeltaNet layers, pushes throughput from 109 to 176 tokens per second on a single H100 in FP8.

The gains carry a cost. "Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840)," the repository notes in its limitations section. On rule-structure and contrastive-policy tasks, though, Jeeves hits 1.000 against Jev's 0.885 and 0.963. It lands alongside NVIDIA's KDA Agent in a busy day of open-source releases.

Shared on Bluesky by 1 AI expert

  • Amy Hoy @amyhoy.bsky.social amplified

    @paulmwatson.com

    Jeeves was inspired by Kev which was inspired by Jev. Models all the way down. No moat. Hope you didn't invest. github.com/PostHog/jeeves

    View on Bluesky →