Hugging Face Paper APO Clusters Users by Preference Bottlenecks for 20-Shot LLM Personalization
Summary
A UW/HKUST paper introduces APO (Approximate Pareto Optimality), a two-stage framework that clusters users by shared 'bottleneck' preference objectives, then meta-trains cluster-specific initializations so new users adapt their LLM aligner from just 20 samples. On Fed-ChatbotPA with Llama-3.2-3B-Instruct APO lifts weighted score to 0.83 vs 0.78 for the best baseline, with up to +17.1% hypervolume gains on UltraFeedback. Uses LoRA (r=8) with DPO loss and 4-bit NF4 quantization.
Originally reported by huggingface.co
Read the original article →Original headline: Hugging Face Paper APO Clusters Users by Preference Bottlenecks for 20-Shot LLM Personalization