The Trickle-down Impact of Reward Inconsistency on RLHF
Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu
摘要
A standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generation. A notable subject that is understudied is the (in-)consistency of RMs -whether they can recognize the semantic changes to different prompts and appropriately adapt its reward assignments -and its impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency:
(1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? ( 3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose CONTRAST INSTRUCTIONS -a benchmarking strategy for the consistency of RM. Each example in CONTRAST INSTRUCTIONS features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on CONTRAST INSTRUCTIONS compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques CONVEXDA and REWARDFU-SION, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.
- Most of the work done while Lingfeng and Sihao were interns at the Tencent AI Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng 等ICLR 2024 · 被引用 299 次
- ARGS: Alignment as Reward-Guided SearchMaxim Khanov, Jirayu Burapacheep, Yixuan LiICLR 2024 · 被引用 101 次
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- Transforming and Combining Rewards for Aligning Large Language ModelsZihao Wang, Chirag Nagpal, Jonathan Berant, Jacob Eisenstein 等ICML 2024 · 被引用 31 次
- Robust Reward Modeling via Causal RubricsPragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil 等ICLR 2026 · 被引用 20 次
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
相关 Paper
- T-REG: Preference Optimization with Token-Level Reward RegularizationWenxuan Zhou, Shujian Zhang, Lingxiao Zhao, Tao MengACL 2025 · 被引用 11 次
- Mitigating Length Bias in RLHF Through a Causal LensHyeonji Kim, Sujeong Oh, Sanghack LeeAAAI 2026 · 被引用 3 次
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang 等NeurIPS 2023 · 被引用 515 次
- Explainable Reinforcement Learning from Human Feedback to Improve AlignmentShicheng Liu, Siyuan Xu, Wenjie Qiu, Hangfan Zhang 等NeurIPS 2025 · 被引用 2 次
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsIlgee Hong, Changlong Yu, Liang Qiu, Weixiang Yan 等NeurIPS 2025 · 被引用 15 次
