'Sharpening Tax' paper: RL post-training cuts pass@K coverage
TL;DR
- The paper tests 14 base/post-trained model pairs across four families and three agentic benchmarks, 42 cases in total.
- Despite worse pass@1, pre-trained LLMs with a light inference harness often beat their post-trained counterparts on pass@K.
- The authors propose posterior-tempered group sampling, a Bayesian sampler that pays a smaller tax than fixed-temperature baselines.
Sharpening Tax in Post-Training, a new arxiv paper, reports that reinforcement-learning post-training of large language models buys higher single-shot accuracy at the price of solution coverage. Base models, sampled many times, can out-cover their RLHF-polished successors on agentic tasks.
The authors, led by Changdae Oh, test the claim across 14 base/post-trained pairs drawn from four model families and three agentic benchmarks, 42 cases in total. The paper reports that "the tax is prevalent in most settings." The mechanism, they argue, is that "post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage." Pre-trained LLMs, "equipped with a light inference harness," they write, "often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget."
As a fix, the paper proposes posterior-tempered group sampling, a Bayesian sampler that adapts temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS "pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy." The abstract does not publish per-benchmark numbers; it reports the tax as prevalent across the 42 cases and defers details to the body.
Originally reported by paper
Read the original article →Original headline: RL Post-Training Kills Multi-Sample Coverage in 36 of 42 Tested Model-Benchmark Pairs