PersonTTS discovers per-user inference policies for LLM reasoning
TL;DR
- PersonTTS treats test-time scaling as finding controllers that jointly satisfy a user's accuracy, latency and inference-cost requirements, rather than one axis at a time.
- The framework reuses prior search experience across users via requirement-matched controller initialization and source-distilled procedural guidance.
- On AIME and HMMT, the authors say it beats strong TTS baselines while cutting discovery-agent time and cost under equal evaluation budgets.
A paper posted to arxiv on October 7, 2026 reframes test-time scaling as a joint problem: pick an inference controller that satisfies a user's accuracy, latency and cost requirements all at once, instead of optimizing one axis at a time. The team, led by Xinglin Wang, calls the framework PersonTTS and publishes it as From Pareto to Preference.
The setup is a flat observation about how users actually ask for inference: "users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers." Each new user profile would normally mean a fresh policy-discovery run. PersonTTS tries to avoid that by reusing prior searches "through requirement-matched controller initialization and source-distilled procedural guidance," while still evaluating every candidate against the target profile.
The reported result, on the AIME and HMMT math benchmarks, is that PersonTTS "substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems." The authors add that under the same candidate-evaluation budget, cross-user experience reuse "further improves policy quality while substantially reducing discovery-agent time and cost."
The abstract puts no numbers on either "substantially," and names only two math-competition benchmarks. Two researchers we track in our Who's Who directory circulated the preprint the same week it landed.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery