Memory Gym: Partially Observable Challenges to Memory-Based Agents
Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss
Abstract
Memory Gym is a novel benchmark for challenging Deep Reinforcement Learning agents to memorize events across long sequences, be robust to noise, and generalize. It consists of the partially observable 2D and discrete control environments Mortar Mayhem, Mystery Path, and Searing Spotlights. These environments are believed to be unsolvable by memory-less agents because they feature strong dependencies on memory and frequent agent-memory interactions. Empirical results based on Proximal Policy Optimization (PPO) and Gated Recurrent Unit (GRU) underline the strong memory dependency of the contributed environments. The hardness of these environments can be smoothly scaled, while different levels of difficulty (some of them unsolved yet) emerge for Mortar Mayhem and Mystery Path. Surprisingly, Searing Spotlights poses a tremendous challenge to GRU-PPO, which remains an open puzzle. Even though therandomly moving spotlights reveal parts of the environment’s ground truth, environmental ablations hint that these pose a severe perturbation to agents that leverage recurrent model architectures as their memory. Source Code: https://github.com/MarcoMeter/drl-memory-gym/
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a0f88834-95dc-44ae-aa52-e01508169226Cited by top-tier papers7
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure et al.ICLR 2024 · 114 citations
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 53 citations
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte et al.ICML 2023 · 21 citations
- QuBE: Question-based Belief Enhancement for Agentic LLM ReasoningMinsoo Kim, Jongyoon Kim, Jihyuk Kim, Seung-won HwangEMNLP 2024 · 2 citations
Related papers
- POPGym: Benchmarking Partially Observable Reinforcement LearningSteven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki et al.ICLR 2023 · 5 citations
- Evaluating Long-Term Memory in 3D MazesJurgis Pasukonis, Timothy P. Lillicrap, Danijar HafnerICLR 2023
- Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement LearningHung Le, Dung Nguyen, Kien Do, Sunil Gupta et al.ICLR 2025
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningEgor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. PanovICLR 2026 · 43 citations
- Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksZexue He, Yu Wang, Churan Zhi, Yuanzhe Hu et al.ICML 2026 · 2 citations
