Goal Misgeneralization in Deep Reinforcement Learning
Lauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau, David Krueger
Abstract
We study goal misgeneralization, a type of out-of-distribution generalization failure in reinforcement learning (RL). Goal misgeneralization failures occur when an RL agent retains its capabilities out-of-distribution yet pursues the wrong goal. For instance, an agent might continue to competently avoid obstacles, but navigate to the wrong place. In contrast, previous works have typically focused on capability generalization failures, where an agent fails to do anything sensible at test time. We formalize this distinction between capability and goal generalization, provide the first empirical demonstrations of goal misgeneralization, and present a partial characterization of its causes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bac89429-9fb3-4fc5-9218-c8813f96a443Cited by top-tier papers30
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba et al.NeurIPS 2023 · 123 citations
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer et al.NeurIPS 2024 · 32 citations
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin et al.ICLR 2026 · 28 citations
Builds on5
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- Optimal Policies Tend To Seek PowerAlexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch et al.NeurIPS 2021 · 111 citations
Related papers
- Horizon Generalization in Reinforcement LearningVivek Myers, Catherine Ji, Benjamin EysenbachICLR 2025
- Learning Domain Invariant Representations in Goal-conditioned Block MDPsBeining Han, Chongyi Zheng, Harris Chan, Keiran Paster et al.NeurIPS 2021 · 20 citations
- Generalization to New Actions in Reinforcement LearningAyush Jain, Andrew Szot, Joseph J. LimICML 2020 · 39 citations
- Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution TrajectoriesLi-Cheng Lan, Huan Zhang, Cho-Jui HsiehICLR 2023 · 3 citations
- Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement LearningRiashat Islam, Hongyu Zang, Anirudh Goyal, Alex Lamb et al.NeurIPS 2022 · 10 citations
