Found first: a primary source the press has not covered yet.
Researchers from Meta Superintelligence Labs, the University of Wisconsin-Madison, Stanford, and NYU show that RL post-training systematically reduces sample diversity in LLMs, a property they call the "sharpening tax." Across 14 base/post-trained model pairs and 3 agentic benchmarks, 36 of 42 model-benchmark combinations showed post-training cutting coverage at scale, even as pass@1 improved. The paper is at arxiv.org/abs/2610.01509.
What the source says
The sharpening tax is defined as base pass@K minus post-trained pass@K, measuring the reduction in test-time scalability that accompanies improved single-shot accuracy. On WebShop with Gemma-4-31B, RL post-training pushed the always-fail rate from 12.4% to 44.0% and the always-pass rate from 0% to 26%, while the middle group of tasks that sometimes passed shrank from 87.6% to 30%. Pass@128 fell from approximately 85% to 56%. The bimodalization pattern held across 36 of 42 combinations spanning Gemma-4 and Qwen2.5 families, among others across 4 model families total. The proposed mitigation, posterior-tempered group sampling (PTGS), estimates difficulty per prompt and adjusts sampling temperature accordingly, achieving simultaneous improvement in pass@1 and pass@K in tested configurations.
Why it matters
RL post-training is the standard path from base model to deployed model. Inference-time scaling depends on running a model repeatedly and selecting the best result, which requires the model to cover diverse solutions across samples. This paper shows post-training degrades that property for most tested combinations, meaning the base model often has better coverage under the same sampling budget. The PTGS mitigation is tested on Qwen2.5-7B-Instruct in constrained environments (Sokoban, FrozenLake); whether the gain holds across the full range of agentic tasks at larger scales is not yet established in this paper.