PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment
Daiwei Chen, Yi Chen, Aniket Rege, Zhi Wang, Ramya Korlakai Vinayak
Abstract
Foundation models trained on internet-scale data benefit from extensive alignment to human preferences before deployment. However, existing methods typically assume a homogeneous preference shared by all individuals, overlooking the diversity inherent in human values. In this work, we propose a general reward modeling framework for pluralistic alignment (PAL), which incorporates diverse preferences from the ground up. PAL has a modular design that leverages commonalities across users while catering to individual personalization, enabling efficient few-shot localization of preferences for new users. Extensive empirical evaluation demonstrates that PAL matches or outperforms state-of-the-art methods on both text-to-text and text-to-image tasks: on Reddit TL;DR Summary, PAL is 1.7% more accurate for seen users and 36% more accurate for unseen users compared to the previous best method, with 100× less parameters. On Pick-a-Pic v2, PAL is 2.5% more accurate than the best method with 156× fewer learned parameters. Finally, we provide theoretical analysis for generalization of rewards learned via PAL showcasing the reduction in number of samples needed per user. Our code is publicly available at https://github.com/RamyaLab/pluralistic-alignment. Figure 1: (a) Using preference data, the PAL framework learns a personalized reward model for each user i, r (i) θ (•), which captures the user's preference for any output x given context xc. (b) PAL models the common perceptions of similarity across users through a shared representation f (•), and represents the individual aspects of preferences via either a preference point a (i) in PAL-A or a preference mapping z (i) (xc) in PAL-B. In particular, we assume a low-rank structure with K prototypical preference points or preference mappings; see Section 2 for details. (c) PAL enables efficient few-shot preference learning for a new user-only a Kdimensional weight vector is learned. This reduces computational cost as well as data needed for generalization.
where h can be any valid link function. The key idea here is that the larger the difference in distances between the alternates to the ideal point, the easier it is to choose between them, and hence the 1 These representations can be taken from penultimate layer(s) of a foundation model. While we use the same D for the prompt and output for simplicity, this can easily be extended to different dimensional spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6db5e0a6-7e1d-4e55-947e-ec55d7761433Cited by top-tier papers8
- Enhancing Personalized Multi-Turn Dialogue with Curiosity RewardYanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani et al.NeurIPS 2025 · 32 citations
- Sparta Alignment: Collectively Aligning Multiple Language Models through CombatYuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett et al.NeurIPS 2025 · 8 citations
- Inference-Time Personalized Alignment with a Few User Preference QueriesVictor-Alexandru Padurean, Parameswaran Kamalaruban, Nachiket Kotalwar, Alkis Gotovos et al.NeurIPS 2025 · 5 citations
- P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic ChecklistKwangwook Seo, Dongha LeeACL 2026 · 2 citations
- PICACO: Pluralistic In-Context Value Alignment via Total Correlation OptimizationHan Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao et al.ICML 2026 · 2 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
Related papers
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
- PAD: Personalized Alignment of LLMs at Decoding-timeRuizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai et al.ICLR 2025
- Beyond Bradley-Terry Models: A General Preference Model for Language Model AlignmentYifan Zhang, Ge Zhang, Yue Wu, Kangping Xu et al.ICML 2025
- One Adapts to Any: Meta Reward Modeling for Personalized LLM AlignmentHongru Cai, Yongqi Li, Tiezheng Yu, Fengbin Zhu et al.SIGIR 2026
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference LearningJingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun et al.EMNLP 2025 · 2 citations
