Interpreting Emergent Planning in Model-Free Reinforcement Learning
Thomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso, David Krueger
摘要
We present the first mechanistic evidence that model-free reinforcement learning agents can learn to plan. This is achieved by applying a methodology based on concept-based interpretability to a model-free agent in Sokoban -a commonly used benchmark for studying planning. Specifically, we demonstrate that DRC, a generic model-free agent introduced by Guez et al. ( 2019 ), uses learned concept representations to internally formulate plans that both predict the long-term effects of actions on the environment and influence action selection. Our methodology involves: (1) probing for planning-relevant concepts, (2) investigating plan formation within the agent's representations, and (3) verifying that discovered plans (in the agent's representations) have a causal effect on the agent's behavior through interventions. We also show that the emergence of these plans coincides with the emergence of a planning-like property: the ability to benefit from additional test-time compute. Finally, we perform a qualitative analysis of the planning algorithm learned by the agent and discover a strong resemblance to parallelized bidirectional search. Our findings advance understanding of the internal mechanisms underlying planning behavior in agents, which is important given the recent trend of emergent planning and reasoning capabilities in LLMs through RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen 等NeurIPS 2024 · 被引用 42 次
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths 等ICLR 2026 · 被引用 7 次
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 被引用 6 次
- Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended EnvironmentsRiley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua 等NeurIPS 2025 · 被引用 6 次
- Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNNMohammad Taufeeque, Aaron David Tucker, Adam Gleave, Adrià Garriga-AlonsoICLR 2026 · 被引用 5 次
它引用的顶会 Paper13
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 被引用 108 次
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha 等ICLR 2020 · 被引用 99 次
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 被引用 78 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
相关 Paper
- Thinker: Learning to Plan and ActStephen Chung, Ivan Anokhin, David KruegerNeurIPS 2023 · 被引用 18 次
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus 等NeurIPS 2022 · 被引用 18 次
- Learning to Search and Searching to Learn for Generalization in PlanningMichael Aichmüller, Yannik Hesse, Hector GeffnerICML 2026
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche 等ICLR 2024 · 被引用 5 次
- SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsMalte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven BehnkeICML 2025
