huggingface.co web signal

'From Pareto to Preference' Paper Amortizes Personalized Test-Time Scaling for LLM Agents

Agents Generative AI ai-business

Summary

Researchers from Beijing Institute of Technology and Xiaohongshu propose amortized agentic policy discovery, which trains a policy to pick decoding strategies conditioned on individual user preferences rather than optimizing to a shared Pareto frontier. The paper argues the approach lets a single model deliver personalized reasoning behavior at test time without per-user fine-tuning.