From 1, 000, 000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment
Jia-Nan Li, Jian Guan, Songhao Wu, Wei Wu, Rui Yan
Abstract
Large language models (LLMs) have traditionally been aligned through one-sizefits-all approaches that assume uniform human preferences, fundamentally overlooking the diversity in user values and needs. This paper introduces a comprehensive framework for scalable personalized alignment of LLMs. We establish a systematic preference space characterizing psychological and behavioral dimensions, alongside diverse persona representations for robust preference inference in real-world scenarios. Building upon this foundation, we introduce ALIGNX, a large-scale dataset of over 1.3 million personalized preference examples, and develop two complementary alignment approaches: in-context alignment directly conditioning on persona representations and preference-bridged alignment modeling intermediate preference distributions. Extensive experiments demonstrate substantial improvements over existing methods, with an average 17.06% accuracy gain across four benchmarks while exhibiting a strong adaptation capability to novel preferences, robustness to limited user data, and precise preference controllability. These results validate our approach toward user-adaptive AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bd2d7d8-e2f6-464e-b57b-23042b7c4aefCited by top-tier papers5
- PersonaVLM: Long-Term Personalized Multimodal LLMsChang Nie, Chaoyou Fu, Yifan Zhang, Haihua Yang et al.CVPR 2026 · 11 citations
- T-POP: Test-Time Personalization with Online Preference FeedbackZikun Qu, Min Zhang, Mingze Kong, Xiang Li et al.ICML 2026 · 4 citations
- Memory OS of AI AgentJiazheng Kang, Mingming Ji, Zhe Zhao, Ting BaiEMNLP 2025 · 4 citations
- CoPL: Collaborative Preference Learning for Personalizing LLMsYoungbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park et al.EMNLP 2025
- PAMDP: Interact to Persona Alignment via a Partially Observable Markov Decision ProcessZhe Yang, Yi Huang, Si Chen, Xiaoting Wu et al.ICLR 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Personality Alignment of Large Language ModelsMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangICLR 2025
- PersonalLLM: Tailoring LLMs to Individual PreferencesThomas P. Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li et al.ICLR 2025
- Align³GR: Unified Multi-Level Alignment for LLM-based Generative RecommendationWencai Ye, Mingjie Sun, Shuhang Chen, Wenjin Wu et al.AAAI 2026 · 2 citations
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith et al.ICLR 2026 · 41 citations
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 8 citations
