huggingface.co via Reddit

Qwen3.8-27B-pi Matches Base xhigh Accuracy With 41% Fewer Tokens

TL;DR

  • On Terminal-Bench 2.1 (89 tasks), Qwen3.8-27B-pi at medium reasoning hits 75.28%, matching the base model's xhigh score with ~41% fewer output tokens.
  • GPQA Diamond xhigh pass rate rises from 80.93% to 86.36%; SciCode xhigh solves 163/337 subproblems (up from 153) with ~23% fewer tokens per solution.
  • The model ships as BF16 (53.81 GB), FP8 (30.39 GB), and 17 GGUF variants from 9 to 29 GB; 4-bit-and-above quants stay within ~1% of BF16 perplexity.

Qwen3.8-27B-pi, a fine-tuned edition of Qwen3.8-27B for the Pi coding agent, matches its base model's xhigh pass rate on Terminal-Bench 2.1 while using roughly 41% fewer output tokens. On the 89-task suite, the tuned model at medium reasoning hits 75.28%, the same score the base needs xhigh reasoning to reach. On GPQA Diamond, xhigh pass rate lifts from 80.93% to 86.36%; on SciCode, xhigh solves 163 of 337 subproblems with about 23% fewer tokens per solved subproblem than the base model.

The fix targets a specific failure mode. In a Hugging Face writeup, author bytkim describes how base Qwen3.8-27B "violated effort ordering": low reasoning used roughly 1.37× more tokens than medium while passing fewer tasks, so raising the effort slider could actually lower quality and raise cost. The tune runs supervised fine-tuning followed by GRPO reinforcement learning with a "Success-Conditioned, Effort-Ordered Reward": failed rollouts earn zero, xhigh gets no length pressure, medium is compared to xhigh's median successful reasoning budget on the same task, and low is compared to whichever higher level solved it.

Two methodology notes carry the judgment. "Validation loss was not a good predictor of agent success," the post reports. The final SFT checkpoint was chosen by running each candidate as an agent, and picking C5 (five steps in) beat picking a later, lower-loss C90 by nearly seven pass points at medium (67.42% vs. 60.67% of 89 tasks). And "Pass/fail rewards waste your best groups": success-conditioned references still produce a learning signal even when every rollout in a group succeeds.

The model ships as BF16 (53.81 GB), FP8 block-wise E4M3 (30.39 GB), and 17 GGUF variants from 9 to 29 GB, with quants at 4-bit and above staying within about 1% of BF16 perplexity. It lands in a busy stretch of open-weights coding releases; our open-source tracker has logged 305 stories over the past 90 days. The author flags the limits plainly: a single training seed, no reward or hyperparameter ablations, and Terminal-Bench 2.1 near saturation at 70–80% pass rates, which leaves limited room to separate models.