PersonTTS Tailors LLM Test-Time Scaling to Per-User Limits
TL;DR
- PersonTTS reframes test-time scaling as finding a controller that jointly meets a user's accuracy, latency, and inference-cost ceilings, not a shared Pareto frontier.
- Experiments span AIME, HMMT and six Qwen3 models at 0.6B through 32B parameters, with 100 source and 20 target profiles per benchmark.
- A frozen 'Guide' distilled from past discovery histories cuts agent-call time about 46% and cost about 36% versus running policy discovery from scratch.
A team from Beijing Institute of Technology and Xiaohongshu has reframed test-time scaling for LLM reasoning as a per-user problem, arguing that chasing a single accuracy-cost or accuracy-latency Pareto frontier misses what real users actually ask for.
The paper, posted to Hugging Face papers, poses personalized TTS as discovering an executable controller that maximizes a 'joint satisfaction rate,' the fraction of evaluation seeds that meet a user's accuracy, latency, and cost ceilings simultaneously. The authors note that running fresh discovery for every new profile is expensive, citing '$39.9 and 160 minutes for a complete policy-discovery run' from the prior AutoTTS system they build on.
PersonTTS amortizes that by retrieving a controller from a source profile with similar requirements, re-evaluating it under the target user, and feeding a frozen 'Guide' distilled from past discovery histories into subsequent revisions. Experiments cover AIME and HMMT with six Qwen3 models at 0.6B, 1.7B, 4B, 8B, 14B, and 32B parameters, 100 source profiles and 20 target profiles per benchmark, and 128 pre-sampled reasoning trajectories per problem-model pair.
The paper reports that Guide-only discovery cuts agent-call time by roughly 46% and cost by about 36% across the two benchmarks, while warm-start retrieval 'mainly shortens elapsed time and can slightly increase cost.' The headline quality gains, though, come mostly from the personalized objective itself: the no-reuse variant 'already preserves most of this advantage,' and enlarging the experience bank helps on AIME held-out problems but 'reverse on HMMT.'
Code and data are released at github.com/WangXinglin/PersonTTS.
Originally reported by huggingface.co
Read the original article →Original headline: 'From Pareto to Preference' Paper Amortizes Personalized Test-Time Scaling for LLM Agents