On the Robustness of Reward Models for Language Model Alignment
Jiwoo Hong, Noah Lee, Eunki Kim, Guijin Son, Woojin Chung, Aman Gupta, Shao Tang, James Thorne
摘要
The Bradley-Terry (BT) model is widely practiced in reward modeling for reinforcement learning with human feedback (RLHF). Despite its effectiveness, reward models (RMs) trained with BT model loss are prone to over-optimization, losing generalizability to unseen input distributions. In this paper, we study the cause of over-optimization in RM training and its downstream effects on the RLHF procedure, accentuating the importance of distributional robustness of RMs in unseen data. First, we show that the excessive dispersion of hidden state norms is the main source of over-optimization. Then, we propose batch-wise sum-to-zero regularization (BSR) to enforce zero-centered reward sum per batch, constraining the rewards with extreme magnitudes. We assess the impact of BSR in improving robustness in RMs through four scenarios of over-optimization, where BSR consistently manifests better robustness. Subsequently, we compare the plain BT model and BSR on RLHF training and empirically show that robust RMs better align the policy to the gold preference model. Finally, we apply BSR to high-quality data and models, which surpasses state-of-theart RMs in the 8B scale by adding more than 5% in complex preference prediction tasks. By conducting RLOO training with 8B RMs, Al-pacaEval 2.0 reduces generation length by 40% while adding a 7% increase in win rate, further highlighting that robustness in RMs induces robustness in RLHF training. We release the code, data, and models: https://github.com/ LinkedIn-XFACT/RM-Robustness .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Why is Your Language Model a Poor Implicit Reward Model?Noam Razin, Yong Lin, Jiarui Yao, Sanjeev AroraICLR 2026 · 被引用 8 次
- Dual-Difficulty Curriculum Learning for Direct Preference OptimizationMengyang Li, Haozhan Geng, Zhong Zhang, Shuang LiuKDD 2026 · 被引用 5 次
- Evaluating and Improving Cultural Awareness of Reward Models for LLM AlignmentHongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang 等ICLR 2026 · 被引用 4 次
- Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level MathGuijin Son, Donghun Yang, Hitesh Patel, Hyunwoo Ko 等ICML 2026 · 被引用 1 次
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward ModelsTong Xie, Ching-Yuan Bai, Yuanhao Ban, Yunqi Hong 等ICML 2026
它引用的顶会 Paper31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
相关 Paper
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang 等NeurIPS 2024 · 被引用 157 次
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 被引用 11 次
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward ModelingZhibin Duan, Guowei Rong, Zhuo Li, Bo Chen 等ICML 2026 · 被引用 4 次
- Towards Cost-Effective Reward Guided Text GenerationAhmad Rashid, Ruotian Wu, Rongqi Fan, Hongliang Li 等ICML 2025
- Beyond Bradley-Terry Models: A General Preference Model for Language Model AlignmentYifan Zhang, Ge Zhang, Yue Wu, Kangping Xu 等ICML 2025
