PRISM: Probabilistic Reward Model with Inherent Structural Modeling
Yuhang Zhou, Yixin Cao, Yuchen Ni, Shihan Dou, Xutian Chen, Ge Zhang, Xiang Liu, Guangnan Ye
摘要
Standard evaluators, such as reward models, compress diverse human judgments into a single scalar, conflating valid Subjective Preference with Epistemic Uncertainty. This structural mismatch often leads to brittle alignment and reward hacking. To address this, we pro-pose PRISM which reinterprets reward evaluation as a conditional distribution parameterized by a Mixture of Gaussians(MOG). PRISM structurally disentangles these factors: distinct Gaussian experts emerge to capture conflicting preference dimensions, while their variance estimates quantify uncertainty, acting as a dynamic reliability gate during optimization. We introduce a two-stage training strategy to learn these disentangled representations from scalable pairwise comparisons without requiring massive fine-grained annotations. Empirical results show that PRISM significantly outperforms scalar baselines in both accuracy and generalization. Furthermore, in downstream Rubric-based Reinforcement Learning, PRISM effectively mitigates reward hacking, yielding policies that are more robust and resilient to distribution shifts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang 等NeurIPS 2024 · 被引用 157 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong 等NeurIPS 2024 · 被引用 60 次
相关 Paper
- Rectifying Shortcut Behaviors in Preference-based Reward LearningWenqian Ye, Guangtao Zheng, Aidong ZhangNeurIPS 2025 · 被引用 6 次
- DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-RewardXiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding 等ACL 2026
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward ModelingZhibin Duan, Guowei Rong, Zhuo Li, Bo Chen 等ICML 2026 · 被引用 4 次
- One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward ModelsDaniel Fein, Max Lamparth, Violet Xiang, Mykel Kochenderfer 等ICML 2026
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi 等ICML 2024 · 被引用 145 次
