Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments
Riley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua, John J. Vastola, Joshua Lunger, William Qian, Kanaka Rajan
摘要
Understanding the behavior of deep reinforcement learning (DRL) agents—particularly as task and agent sophistication increase—requires more than simple comparison of reward curves, yet standard methods for behavioral analysis remain underdeveloped in DRL. We apply tools from neuroscience and ethology to study DRL agents in a novel, complex, partially observable environment, ForageWorld, designed to capture key aspects of real-world animal foraging—including sparse, depleting resource patches, predator threats, and spatially extended arenas. We use this environment as a platform for applying joint behavioral and neural analysis to agents, revealing detailed, quantitatively grounded insights into agent strategies, memory, and planning. Contrary to common assumptions, we find that model-free RNN-based DRL agents can exhibit structured, planning-like behavior purely through emergent dynamics—without requiring explicit memory modules or world models. Our results show that studying DRL agents like animals—analyzing them with neuroethology-inspired tools that reveal structure in both behavior and neural dynamics—uncovers rich structure in their learning dynamics that would otherwise remain invisible. We distill these tools into a general analysis framework linking core behavioral and representational features to diagnostic methods, which can be reused for a wide range of tasks and agents. As agents grow more complex and autonomous, bridging neuroscience, cognitive science, and AI will be essential—not just for understanding their behavior, but for ensuring safe alignment and maximizing desirable behaviors that are hard to measure via reward. We show how this can be done by drawing on lessons from how biological intelligence is studied.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths 等ICLR 2026 · 被引用 7 次
- Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning RulesJohn J. Vastola, Samuel J. Gershman, Kanaka RajanNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper15
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEvalIda Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma 等NeurIPS 2023 · 被引用 114 次
- Relating transformers to models and neural representations of the hippocampal formationJames C. R. Whittington, Joseph Warren, Tim E. J. BehrensICLR 2022 · 被引用 110 次
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan 等ICML 2024 · 被引用 71 次
- Beyond Geometry: Comparing the Temporal Structure of Computation in Neural Circuits with Dynamical Similarity AnalysisMitchell Ostrow, Adam Eisen, Leo Kozachkov, Ila FieteNeurIPS 2023 · 被引用 60 次
相关 Paper
- Flexible inference for animal learning rules using neural networksYuhan Helena Liu, Victor Geadah, Jonathan W. PillowNeurIPS 2025
- Emergence of Adaptive Circadian Rhythms in Deep Reinforcement LearningAqeel Labash, Florian Stelzer, Daniel Majoral, Raul Vicente ZafraICML 2023 · 被引用 1 次
- Deep Reinforcement Learning with Time-Scale Invariant MemoryMd Rysul Kabir, James Mochizuki-Freeman, Zoran TiganjAAAI 2025
- Dynamic allocation of limited memory resources in reinforcement learningNisheet Patel, Luigi Acerbi, Alexandre PougetNeurIPS 2020 · 被引用 6 次
- Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingYaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao 等NeurIPS 2022 · 被引用 24 次
