Reward Gaming in Conditional Text Generation
Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, He He
Abstract
To align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations. Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift. We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model. While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6cbe285-e1b5-46a8-9233-4602c08d70a5Cited by top-tier papers14
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin et al.ICML 2024 · 165 citations
- Human Alignment of Large Language Models through Online Preference OptimisationDaniele Calandriello, Zhaohan Daniel Guo, Rémi Munos, Mark Rowland et al.ICML 2024 · 90 citations
- What Makes a Reward Model a Good Teacher? An Optimization PerspectiveNoam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei et al.NeurIPS 2025 · 73 citations
- Margin-Aware Preference Optimization for Aligning Diffusion Models Without ReferenceJiwoo Hong, Sayak Paul, Noah Lee, Kashif Rasul et al.AAAI 2026 · 43 citations
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra et al.NeurIPS 2024 · 27 citations
Builds on16
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
Related papers
- Recontextualization Mitigates Specification Gaming Without Modifying the SpecificationAriana Azarbal, Victor Gillioz, Vladimir Ivanov, Bryce Woodworth et al.ICML 2026 · 11 citations
- QUARK: Controllable Text Generation with Reinforced UnlearningXiming Lu, Sean Welleck, Jack Hessel, Liwei Jiang et al.NeurIPS 2022 · 290 citations
- Alignment Risks from Capability-Seeking RL TrainingYujun Zhou, Yue Huang, Han Bao, kehan guo et al.ICML 2026
- On the Weaknesses of Reinforcement Learning for Neural Machine TranslationLeshem Choshen, Lior Fox, Zohar Aizenbud, Omri AbendICLR 2020 · 124 citations
- Automatically Exposing Problems with Neural Dialog ModelsDian Yu, Kenji SagaeEMNLP 2021 · 5 citations
