Mitigating Reward Overoptimization via Lightweight Uncertainty Estimation
Xiaoying Zhang, Jean-Francois Ton, Wei Shen, Hongning Wang, Yang Liu
摘要
Reinforcement Learning from Human Feedback (RLHF) has been pivotal in aligning Large Language Models with human values but often suffers from overopti-mization due to its reliance on a proxy reward model. To mitigate this limitation, we first propose a lightweight uncertainty quantification method that assesses the reliability of the proxy reward using only the last layer embeddings of the reward model. Enabled by this efficient uncertainty quantification method, we formulate A DV PO, a distributionally robust optimization procedure to tackle the reward overoptimization problem in RLHF. Through extensive experiments on the An-thropic HH and TL;DR summarization datasets, we verify the effectiveness of A DV PO in mitigating the overoptimization problem, resulting in enhanced RLHF performance as evaluated through human-assisted evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Ask a Strong LLM Judge when Your Reward Model is UncertainZhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu 等NeurIPS 2025 · 被引用 14 次
- Bradley-Terry and Multi-Objective Reward Modeling Are ComplementaryZhiwei Zhang, Hui Liu, Xiaomin Li, Zhenwei Dai 等ICLR 2026 · 被引用 8 次
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward ModelingZhibin Duan, Guowei Rong, Zhuo Li, Bo Chen 等ICML 2026 · 被引用 4 次
- On the Robustness of Reward Models for Language Model AlignmentJiwoo Hong, Noah Lee, Eunki Kim, Guijin Son 等ICML 2025
- Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N SamplingJizhou Guo, Zhaomin Wu, Hanchen Yang, Philip S. YuKDD 2026
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
相关 Paper
- Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHFShicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai 等ICLR 2025
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 被引用 208 次
- Distributionally Robust Reinforcement Learning from Human FeedbackDebmalya Mandal, Paulius Sasnauskas, Goran RadanovicICML 2026
- Adaptive Preference Scaling for Reinforcement Learning with Human FeedbackIlgee Hong, Zichong Li, Alexander Bukharin, Yixiao Li 等NeurIPS 2024 · 被引用 23 次
- Reliability-Aware LLM Alignment from Inconsistent Human FeedbackJingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma 等ICML 2026
