Lune

ICLR2025Top-tier venue

PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment

Daiwei Chen, Yi Chen, Aniket Rege, Zhi Wang, Ramya Korlakai Vinayak

2025Year
8Top-tier citations

Abstract

Foundation models trained on internet-scale data benefit from extensive alignment to human preferences before deployment. However, existing methods typically assume a homogeneous preference shared by all individuals, overlooking the diversity inherent in human values. In this work, we propose a general reward modeling framework for pluralistic alignment (PAL), which incorporates diverse preferences from the ground up. PAL has a modular design that leverages commonalities across users while catering to individual personalization, enabling efficient few-shot localization of preferences for new users. Extensive empirical evaluation demonstrates that PAL matches or outperforms state-of-the-art methods on both text-to-text and text-to-image tasks: on Reddit TL;DR Summary, PAL is 1.7% more accurate for seen users and 36% more accurate for unseen users compared to the previous best method, with 100× less parameters. On Pick-a-Pic v2, PAL is 2.5% more accurate than the best method with 156× fewer learned parameters. Finally, we provide theoretical analysis for generalization of rewards learned via PAL showcasing the reduction in number of samples needed per user. Our code is publicly available at https://github.com/RamyaLab/pluralistic-alignment. Figure 1: (a) Using preference data, the PAL framework learns a personalized reward model for each user i, r (i) θ (•), which captures the user's preference for any output x given context xc. (b) PAL models the common perceptions of similarity across users through a shared representation f (•), and represents the individual aspects of preferences via either a preference point a (i) in PAL-A or a preference mapping z (i) (xc) in PAL-B. In particular, we assume a low-rank structure with K prototypical preference points or preference mappings; see Section 2 for details. (c) PAL enables efficient few-shot preference learning for a new user-only a Kdimensional weight vector is learned. This reduces computational cost as well as data needed for generalization.

where h can be any valid link function. The key idea here is that the larger the difference in distances between the alternates to the ideal point, the easier it is to choose between them, and hence the 1 These representations can be taken from penultimate layer(s) of a foundation model. While we use the same D for the prompt and output for simplicity, this can easily be extended to different dimensional spaces.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6db5e0a6-7e1d-4e55-947e-ec55d7761433

Cited by top-tier papers8

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines