Mitigating Reward Overoptimization via Lightweight Uncertainty Estimation
Xiaoying Zhang, Jean-Francois Ton, Wei Shen, Hongning Wang, Yang Liu
Abstract
Reinforcement Learning from Human Feedback (RLHF) has been pivotal in aligning Large Language Models with human values but often suffers from overopti-mization due to its reliance on a proxy reward model. To mitigate this limitation, we first propose a lightweight uncertainty quantification method that assesses the reliability of the proxy reward using only the last layer embeddings of the reward model. Enabled by this efficient uncertainty quantification method, we formulate A DV PO, a distributionally robust optimization procedure to tackle the reward overoptimization problem in RLHF. Through extensive experiments on the An-thropic HH and TL;DR summarization datasets, we verify the effectiveness of A DV PO in mitigating the overoptimization problem, resulting in enhanced RLHF performance as evaluated through human-assisted evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd924e18-3efd-4629-97d0-4e8551ce616cCited by top-tier papers5
- Ask a Strong LLM Judge when Your Reward Model is UncertainZhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu et al.NeurIPS 2025 · 14 citations
- Bradley-Terry and Multi-Objective Reward Modeling Are ComplementaryZhiwei Zhang, Hui Liu, Xiaomin Li, Zhenwei Dai et al.ICLR 2026 · 8 citations
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward ModelingZhibin Duan, Guowei Rong, Zhuo Li, Bo Chen et al.ICML 2026 · 4 citations
- On the Robustness of Reward Models for Language Model AlignmentJiwoo Hong, Noah Lee, Eunki Kim, Guijin Son et al.ICML 2025
- Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N SamplingJizhou Guo, Zhaomin Wu, Hanchen Yang, Philip S. YuKDD 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHFShicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai et al.ICLR 2025
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 208 citations
- Distributionally Robust Reinforcement Learning from Human FeedbackDebmalya Mandal, Paulius Sasnauskas, Goran RadanovicICML 2026
- Adaptive Preference Scaling for Reinforcement Learning with Human FeedbackIlgee Hong, Zichong Li, Alexander Bukharin, Yixiao Li et al.NeurIPS 2024 · 23 citations
- Reliability-Aware LLM Alignment from Inconsistent Human FeedbackJingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma et al.ICML 2026
