Lune

ICLR2025顶会

PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment

Daiwei Chen, Yi Chen, Aniket Rege, Zhi Wang, Ramya Korlakai Vinayak

出版方
2025年份
8顶会引用

摘要

Foundation models trained on internet-scale data benefit from extensive alignment to human preferences before deployment. However, existing methods typically assume a homogeneous preference shared by all individuals, overlooking the diversity inherent in human values. In this work, we propose a general reward modeling framework for pluralistic alignment (PAL), which incorporates diverse preferences from the ground up. PAL has a modular design that leverages commonalities across users while catering to individual personalization, enabling efficient few-shot localization of preferences for new users. Extensive empirical evaluation demonstrates that PAL matches or outperforms state-of-the-art methods on both text-to-text and text-to-image tasks: on Reddit TL;DR Summary, PAL is 1.7% more accurate for seen users and 36% more accurate for unseen users compared to the previous best method, with 100× less parameters. On Pick-a-Pic v2, PAL is 2.5% more accurate than the best method with 156× fewer learned parameters. Finally, we provide theoretical analysis for generalization of rewards learned via PAL showcasing the reduction in number of samples needed per user. Our code is publicly available at https://github.com/RamyaLab/pluralistic-alignment. Figure 1: (a) Using preference data, the PAL framework learns a personalized reward model for each user i, r (i) θ (•), which captures the user's preference for any output x given context xc. (b) PAL models the common perceptions of similarity across users through a shared representation f (•), and represents the individual aspects of preferences via either a preference point a (i) in PAL-A or a preference mapping z (i) (xc) in PAL-B. In particular, we assume a low-rank structure with K prototypical preference points or preference mappings; see Section 2 for details. (c) PAL enables efficient few-shot preference learning for a new user-only a Kdimensional weight vector is learned. This reduces computational cost as well as data needed for generalization.

where h can be any valid link function. The key idea here is that the larger the difference in distances between the alternates to the ideal point, the easier it is to choose between them, and hence the 1 These representations can be taken from penultimate layer(s) of a foundation model. While we use the same D for the prompt and output for simplicity, this can easily be extended to different dimensional spaces.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖