huggingface.co web signal

Hugging Face Paper APO Clusters Users by Preference Bottlenecks for 20-Shot LLM Personalization

Fine-tuning Open Source ai-business

Summary

A UW/HKUST paper introduces APO (Approximate Pareto Optimality), a two-stage framework that clusters users by shared 'bottleneck' preference objectives, then meta-trains cluster-specific initializations so new users adapt their LLM aligner from just 20 samples. On Fed-ChatbotPA with Llama-3.2-3B-Instruct APO lifts weighted score to 0.83 vs 0.78 for the best baseline, with up to +17.1% hypervolume gains on UltraFeedback. Uses LoRA (r=8) with DPO loss and 4-bit NF4 quantization.