Interpreting Emergent Planning in Model-Free Reinforcement Learning
Thomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso, David Krueger
Abstract
We present the first mechanistic evidence that model-free reinforcement learning agents can learn to plan. This is achieved by applying a methodology based on concept-based interpretability to a model-free agent in Sokoban -a commonly used benchmark for studying planning. Specifically, we demonstrate that DRC, a generic model-free agent introduced by Guez et al. ( 2019 ), uses learned concept representations to internally formulate plans that both predict the long-term effects of actions on the environment and influence action selection. Our methodology involves: (1) probing for planning-relevant concepts, (2) investigating plan formation within the agent's representations, and (3) verifying that discovered plans (in the agent's representations) have a causal effect on the agent's behavior through interventions. We also show that the emergence of these plans coincides with the emergence of a planning-like property: the ability to benefit from additional test-time compute. Finally, we perform a qualitative analysis of the planning algorithm learned by the agent and discover a strong resemblance to parallelized bidirectional search. Our findings advance understanding of the internal mechanisms underlying planning behavior in agents, which is important given the recent trend of emergent planning and reasoning capabilities in LLMs through RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d16a4eb-fe72-4145-bf2f-f113ea02537eCited by top-tier papers11
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen et al.NeurIPS 2024 · 42 citations
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths et al.ICLR 2026 · 7 citations
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 6 citations
- Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended EnvironmentsRiley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua et al.NeurIPS 2025 · 6 citations
- Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNNMohammad Taufeeque, Aaron David Tucker, Adam Gleave, Adrià Garriga-AlonsoICLR 2026 · 5 citations
Builds on13
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha et al.ICLR 2020 · 99 citations
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 78 citations
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
Related papers
- Thinker: Learning to Plan and ActStephen Chung, Ivan Anokhin, David KruegerNeurIPS 2023 · 18 citations
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus et al.NeurIPS 2022 · 18 citations
- Learning to Search and Searching to Learn for Generalization in PlanningMichael Aichmüller, Yannik Hesse, Hector GeffnerICML 2026
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche et al.ICLR 2024 · 5 citations
- SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsMalte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven BehnkeICML 2025
