huggingface.co web signal

APO Paper Personalizes LLM Alignment From 20 User Samples

Fine-tuning Open Source ai-business

TL;DR

  • APO adapts a per-user LLM preference aligner from just 20 local examples, after clustering users by shared 'bottleneck' preference objectives.
  • On Fed-ChatbotPA with Llama-3.2-3B-Instruct, the preference-weighted score reaches 0.83 against 0.78 for the best baseline.
  • On UltraFeedback the Pareto hypervolume rises 0.728 to 0.820 (+12.6%), using LoRA rank 8 with DPO loss under 4-bit NF4.

A new preprint proposes personalizing a language model's preference aligner from just 20 user samples, after first grouping users whose gradient updates don't fight each other.

The method, Approximate Pareto Optimality or APO, comes from Yige Yuan at the University of Washington and Zhiqin Yang at the Hong Kong University of Science and Technology, listed as corresponding author. They frame the problem plainly: 'Real-world users exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult.'

The fix is two-stage. First, Bottleneck-Adjustment Clustering partitions clients by their shared bottleneck preference objective, then refines each cluster by adjustment-vector direction, so updates inside a cluster descend together instead of cancelling. Second, that cluster-specific initialization is meta-trained for K=3 adaptation steps against S=20-shot episodes, so a new user can adapt from 20 local examples.

On Fed-ChatbotPA with Llama-3.2-3B-Instruct, the preference-weighted score reaches 0.83 against 0.78 for the best baseline. On UltraFeedback, Pareto hypervolume climbs from 0.728 to 0.820, a +12.6% gain. The paper reports APO 'achieves highest Score on 23/24 model-dataset-cluster combinations,' and highest or tied-highest hypervolume on 22/24.

The recipe is cheap: LoRA with rank 8 and alpha 16, DPO loss (β=0.1), 4-bit NF4 quantization, 200 phase-1 rounds plus 10 outer rounds of meta-training, batch size 8. Baselines compared include RIC, Rewarded Soup, FSPO, DITTO and FedAvg. Experiments cover two datasets and two 3B-parameter base models (Llama-3.2-3B-Instruct and Qwen2.5-3B-Instruct); larger production-scale models are not reported. Source code is 'to be released upon acceptance.'