Investigating Memory in RL with POPGym Arcade
Zekang Wang, Zhe He, Borong Zhang, Edan Toledo, Steven Morad
Abstract
How should we analyze memory in deep RL? We introduce tools for analyzing policies under partial observability and revealing how agents use memory to make decisions. To utilize these tools, we present POPGym Arcade, a collection of Atari-inspired, hardware-accelerated environments sharing a single observation and action space. Each environment provides fully and partially observable variants, enabling counterfactual studies on observability. We find that controlled studies are necessary for fair comparisons and identify a pathology where value functions smear credit over irrelevant history. Using this pathology, we demonstrate how out-ofdistribution scenarios can contaminate memory, perturbing the policy far into the future. Our code is available at https://github.com/bol t-research/popgym-arcade .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d55f386-a62d-42d4-bfba-06221b44f319Builds on21
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
Related papers
- POPGym: Benchmarking Partially Observable Reinforcement LearningSteven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki et al.ICLR 2023 · 5 citations
- Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement LearningHung Le, Dung Nguyen, Kien Do, Sunil Gupta et al.ICLR 2025
- Memory Gym: Partially Observable Challenges to Memory-Based AgentsMarco Pleines, Matthias Pallasch, Frank Zimmer, Mike PreussICLR 2023
- Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksZexue He, Yu Wang, Churan Zhi, Yuanzhe Hu et al.ICML 2026 · 2 citations
- Recurrent Action Transformer with MemoryEgor Cherepanov, Aleksei Staroverov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 14 citations
