'From Pareto to Preference' Paper Amortizes Personalized Test-Time Scaling for LLM Agents
Summary
Researchers from Beijing Institute of Technology and Xiaohongshu propose amortized agentic policy discovery, which trains a policy to pick decoding strategies conditioned on individual user preferences rather than optimizing to a shared Pareto frontier. The paper argues the approach lets a single model deliver personalized reasoning behavior at test time without per-user fine-tuning.
Originally reported by huggingface.co
Read the original article →Original headline: 'From Pareto to Preference' Paper Amortizes Personalized Test-Time Scaling for LLM Agents