The Trickle-down Impact of Reward Inconsistency on RLHF
Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu
Abstract
A standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generation. A notable subject that is understudied is the (in-)consistency of RMs -whether they can recognize the semantic changes to different prompts and appropriately adapt its reward assignments -and its impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency:
(1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? ( 3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose CONTRAST INSTRUCTIONS -a benchmarking strategy for the consistency of RM. Each example in CONTRAST INSTRUCTIONS features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on CONTRAST INSTRUCTIONS compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques CONVEXDA and REWARDFU-SION, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.
- Most of the work done while Lingfeng and Sihao were interns at the Tencent AI Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5368bec4-8d0e-4570-a26b-7feca8e17300Cited by top-tier papers17
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- ARGS: Alignment as Reward-Guided SearchMaxim Khanov, Jirayu Burapacheep, Yixuan LiICLR 2024 · 101 citations
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju et al.NeurIPS 2025 · 38 citations
- Transforming and Combining Rewards for Aligning Large Language ModelsZihao Wang, Chirag Nagpal, Jonathan Berant, Jacob Eisenstein et al.ICML 2024 · 31 citations
- Robust Reward Modeling via Causal RubricsPragya Srivastava, Harman Singh, Rahul Madhavan, Gandharv Patil et al.ICLR 2026 · 20 citations
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
Related papers
- T-REG: Preference Optimization with Token-Level Reward RegularizationWenxuan Zhou, Shujian Zhang, Lingxiao Zhao, Tao MengACL 2025 · 11 citations
- Mitigating Length Bias in RLHF Through a Causal LensHyeonji Kim, Sujeong Oh, Sanghack LeeAAAI 2026 · 3 citations
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang et al.NeurIPS 2023 · 515 citations
- Explainable Reinforcement Learning from Human Feedback to Improve AlignmentShicheng Liu, Siyuan Xu, Wenjie Qiu, Hangfan Zhang et al.NeurIPS 2025 · 2 citations
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsIlgee Hong, Changlong Yu, Liang Qiu, Weixiang Yan et al.NeurIPS 2025 · 15 citations
