DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
Shaofan Liu, Guoqiang Zhang, Shihan Dou, Huiyuan Zheng, Yiming Zhou, Junjie Ye, Shaowen Wang, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang, Xuanjing Huang
摘要
Reward models (RMs) are the surrogate objectives in reinforcement learning from human feedback (RLHF), and their scores directly steer policy optimization. We show that standard RM training is vulnerable in data subsets where response quality depends only weakly on the context: such instances encourage the RM to ignore the context, leading to context neglect and degraded accuracy. To address this failure mode, we propose Distribution-Aware Reward Modeling (DARM), which augments the RM objective with a conditional mutual information regularizer that maximizes context and the predicted reward conditioned on the response. By explicitly preserving the sensitivity of reward signals to the prompting context, DARM reduces over-reliance on response-only features and improves robustness to contextual variation. Extensive experiments across in-distribution and out-of-distribution settings show that DARM trained RMs deliver more accurate and consistent scoring than strong baselines. We further evaluate its downstream impact in RLHF, where DARM produces better aligned policies. We also demonstrate the necessity of each DARM design component and the impact of key parameters on performance through ablation experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
相关 Paper
- Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningShihan Dou, Yan Liu, Enyu Zhou, Songyang Gao 等AAAI 2025 · 被引用 2 次
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang 等NeurIPS 2024 · 被引用 157 次
- Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference ModelingShihan Dou, Jiayi Chen, Chenhao Huang, Feng Chen 等ACL 2025
- Information-Theoretic Reward Decomposition for Generalizable RLHFLiyuan Mao, Haoran Xu, Amy Zhang, Weinan Zhang 等NeurIPS 2025 · 被引用 7 次
- Rethinking Reward Model Evaluation Through the Lens of Reward OveroptimizationSunghwan Kim, Dongjin Kang, Taeyoon Kwon, Hyungjoo Chae 等ACL 2025
