Goal Misgeneralization in Deep Reinforcement Learning
Lauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau, David Krueger
摘要
We study goal misgeneralization, a type of out-of-distribution generalization failure in reinforcement learning (RL). Goal misgeneralization failures occur when an RL agent retains its capabilities out-of-distribution yet pursues the wrong goal. For instance, an agent might continue to competently avoid obstacles, but navigate to the wrong place. In contrast, previous works have typically focused on capability generalization failures, where an agent fails to do anything sensible at test time. We formalize this distinction between capability and goal generalization, provide the first empirical demonstrations of goal misgeneralization, and present a partial characterization of its causes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya 等NeurIPS 2023 · 被引用 295 次
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba 等NeurIPS 2023 · 被引用 123 次
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer 等NeurIPS 2024 · 被引用 32 次
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin 等ICLR 2026 · 被引用 28 次
它引用的顶会 Paper5
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 被引用 293 次
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 被引用 120 次
- Optimal Policies Tend To Seek PowerAlexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch 等NeurIPS 2021 · 被引用 111 次
相关 Paper
- Horizon Generalization in Reinforcement LearningVivek Myers, Catherine Ji, Benjamin EysenbachICLR 2025
- Learning Domain Invariant Representations in Goal-conditioned Block MDPsBeining Han, Chongyi Zheng, Harris Chan, Keiran Paster 等NeurIPS 2021 · 被引用 20 次
- Generalization to New Actions in Reinforcement LearningAyush Jain, Andrew Szot, Joseph J. LimICML 2020 · 被引用 39 次
- Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution TrajectoriesLi-Cheng Lan, Huan Zhang, Cho-Jui HsiehICLR 2023 · 被引用 3 次
- Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement LearningRiashat Islam, Hongyu Zang, Anirudh Goyal, Alex Lamb 等NeurIPS 2022 · 被引用 10 次
