Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
Zhuo Li, Pengyu Cheng, Zhechao Yu, FeifeiTong, Anningzhe Gao, Tsung-Hui Chang, Xiang Wan, erchao.zec, xiaoxi jiang, guanjunjiang
Abstract
Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values. However, RM training data is commonly recognized as low-quality, containing inductive biases that can easily lead to overfitting and reward hacking. For example, more detailed and comprehensive responses are usually human-preferred but with more words, leading response length to become one of the inevitable inductive biases. A limited number of prior RM debiasing approaches either target a single specific type of bias or model the problem with only simple linear correlations, e.g., Pearson coefficients. To mitigate more complex and diverse inductive biases in reward modeling, we introduce a novel information-theoretic debiasing method called Debiasing via Information optimization for RM (DIR). Inspired by the information bottleneck (IB), we maximize the mutual information (MI) between RM scores and human preference pairs, while minimizing the MI between RM outputs and biased attributes of preference inputs. With theoretical justification from information theory, DIR can handle more sophisticated types of biases with non-linear correlations, broadly extending the real-world application scenarios for RM debiasing methods. In experiments, we verify the effectiveness of DIR with three types of inductive biases: response length, sycophancy, and format. We discover that DIR not only effectively mitigates target inductive biases but also enhances RLHF performance across diverse benchmarks, yielding better generalization abilities. The code and training recipes are available at https://github.com/Qwen-Applications/DIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHFKangwen Zhao, Jianfeng Cai, Jinhua Zhu, Ruopei Sun et al.ACL 2026 · 7 citations
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingYuchun Miao, Sen Zhang, Liang Ding, Rong Bao et al.NeurIPS 2024 · 108 citations
- One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward ModelsDaniel Fein, Max Lamparth, Violet Xiang, Mykel Kochenderfer et al.ICML 2026
- ODIN: Disentangled Reward Mitigates Hacking in RLHFLichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia et al.ICML 2024 · 119 citations
- Disentangling Length Bias in Preference Learning via Response-Conditioned ModelingJianfeng Cai, Jinhua Zhu, Ruopei Sun, Yue Wang et al.ICLR 2026 · 6 citations
