Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments
Riley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua, John J. Vastola, Joshua Lunger, William Qian, Kanaka Rajan
Abstract
Understanding the behavior of deep reinforcement learning (DRL) agents—particularly as task and agent sophistication increase—requires more than simple comparison of reward curves, yet standard methods for behavioral analysis remain underdeveloped in DRL. We apply tools from neuroscience and ethology to study DRL agents in a novel, complex, partially observable environment, ForageWorld, designed to capture key aspects of real-world animal foraging—including sparse, depleting resource patches, predator threats, and spatially extended arenas. We use this environment as a platform for applying joint behavioral and neural analysis to agents, revealing detailed, quantitatively grounded insights into agent strategies, memory, and planning. Contrary to common assumptions, we find that model-free RNN-based DRL agents can exhibit structured, planning-like behavior purely through emergent dynamics—without requiring explicit memory modules or world models. Our results show that studying DRL agents like animals—analyzing them with neuroethology-inspired tools that reveal structure in both behavior and neural dynamics—uncovers rich structure in their learning dynamics that would otherwise remain invisible. We distill these tools into a general analysis framework linking core behavioral and representational features to diagnostic methods, which can be reused for a wide range of tasks and agents. As agents grow more complex and autonomous, bridging neuroscience, cognitive science, and AI will be essential—not just for understanding their behavior, but for ensuring safe alignment and maximizing desirable behaviors that are hard to measure via reward. We show how this can be done by drawing on lessons from how biological intelligence is studied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ba0aec4-e7eb-4e7e-9366-6705b3800138Cited by top-tier papers2
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths et al.ICLR 2026 · 7 citations
- Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning RulesJohn J. Vastola, Samuel J. Gershman, Kanaka RajanNeurIPS 2025 · 4 citations
Builds on15
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEvalIda Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma et al.NeurIPS 2023 · 114 citations
- Relating transformers to models and neural representations of the hippocampal formationJames C. R. Whittington, Joseph Warren, Tim E. J. BehrensICLR 2022 · 110 citations
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan et al.ICML 2024 · 71 citations
- Beyond Geometry: Comparing the Temporal Structure of Computation in Neural Circuits with Dynamical Similarity AnalysisMitchell Ostrow, Adam Eisen, Leo Kozachkov, Ila FieteNeurIPS 2023 · 60 citations
Related papers
- Flexible inference for animal learning rules using neural networksYuhan Helena Liu, Victor Geadah, Jonathan W. PillowNeurIPS 2025
- Emergence of Adaptive Circadian Rhythms in Deep Reinforcement LearningAqeel Labash, Florian Stelzer, Daniel Majoral, Raul Vicente ZafraICML 2023 · 1 citation
- Deep Reinforcement Learning with Time-Scale Invariant MemoryMd Rysul Kabir, James Mochizuki-Freeman, Zoran TiganjAAAI 2025
- Dynamic allocation of limited memory resources in reinforcement learningNisheet Patel, Luigi Acerbi, Alexandre PougetNeurIPS 2020 · 6 citations
- Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingYaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao et al.NeurIPS 2022 · 24 citations
