'Sharpening Tax' Paper: RL Post-Training Hurts pass@K Coverage
TL;DR
- The paper studies 14 base/post-trained model pairs across four model families and three agentic benchmarks, 42 cases in total.
- Base LLMs with a light inference harness often beat post-trained versions on pass@K solution coverage despite lower pass@1 accuracy.
- The proposed PTGS Bayesian sampler adapts temperature per prompt difficulty and pays a smaller tax than fixed-temperature baselines.
RL post-training of LLMs comes with a measurable cost on test-time scalability, with base models equipped only with a light inference harness often outscoring their post-trained counterparts on multi-sample coverage. That is the finding of a Hugging Face paper proposing 'Sharpening Tax' as a diagnostic metric.
The authors, led by Changdae Oh, tested the pattern across 14 base/post-trained model pairs from four model families and three agentic benchmarks, 42 cases in all. "Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget," the paper reports, and the tax "is prevalent in most settings."
The mechanism, the authors write, is that "post-training pushes tasks toward two extremes, always solved or never solved," improving sampling efficiency and consistency on the solvable ones at the cost of ever reaching new ones.
Their fix is posterior-tempered group sampling, or PTGS, "a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty." Applied during RL training in two agentic environments, PTGS "pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy." It is one of about 50 fine-tuning papers in our tracker from the last 90 days probing where RL post-training helps and where it hurts.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper 'Sharpening Tax' Shows RL Post-Training Hurts pass@K Coverage, Introduces Posterior-Tempered Sampling Fix