github.com via Hacker News

Firelex ships Jeff, home-trained Jev-compatible decision models

TL;DR

  • The 0.8B Qwen variant returns a decision in about 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max.
  • Across five public benchmarks the 2B Qwen model scores 83.1%, narrowly edging Jev at 83.0%; the 0.8B lands at 79.1%.
  • Training ran entirely on one RTX PRO 6000 workstation GPU, about two hours for the 0.8B and 3.5 hours for the 2B.

Firelex has published Jeff, a trio of small decision models fine-tuned from Qwen3.5 and Gemma 4 that return calibrated probabilities over user-defined options in a single forward pass. Median inference sits at about 22 ms per decision on an RTX PRO 6000 workstation GPU and 28 ms on an Apple M4 Max, per the project's GitHub README.

The three variants, a 0.8B Qwen, a 2B Qwen and a Gemma4-E2B, post combined scores of 79.1%, 83.1% and 81.6% across five public benchmarks, with the 2B narrowly edging Jev's 83.0%. The developer built the whole thing on a single workstation, noting the 0.8B trained in about two hours and the 2B in about 3.5. Synthetic training data came from an open model, and the recipe descends from the MIT-licensed AutoJev.

The scope is explicit. "Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning," the README says, pitching Jeff at support-queue routing, content moderation, voice commands and game AI rather than general assistants. Compatibility gets its own line: "Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev."

It lands the same day Supersonic Labs open-weighted the 144M Julia-1, making small decision models a live thread in our open-source feed.