POPGym: Benchmarking Partially Observable Reinforcement Learning
Steven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki, Amanda Prorok
Abstract
Real world applications of Reinforcement Learning (RL) are often partially observable, thus requiring memory. Despite this, partial observability is still largely ignored by contemporary RL benchmarks and libraries. We introduce Partially Observable Process Gym (POPGym), a two-part library containing (1) a diverse collection of 15 partially observable environments, each with multiple difficulties and (2) implementations of 13 memory model baselines -- the most in a single RL library. Existing partially observable benchmarks tend to fixate on 3D visual navigation, which is computationally expensive and only one type of POMDP. In contrast, POPGym environments are diverse, produce smaller observations, use less memory, and often converge within two hours of training on a consumer-grade GPU. We implement our high-level memory API and memory baselines on top of the popular RLlib framework, providing plug-and-play compatibility with various training algorithms, exploration strategies, and distributed training paradigms. Using POPGym, we execute the largest comparison across RL memory models to date. POPGym is available at https://github.com/proroklab/popgym.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58fc8dea-1fec-4d7e-99fe-7b6fb0049068Cited by top-tier papers20
- Structured State Space Models for In-Context Reinforcement LearningChris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto et al.NeurIPS 2023 · 164 citations
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningEgor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. PanovICLR 2026 · 43 citations
- Mastering Memory Tasks with World ModelsMohammad Reza Samsami, Artem Zholus, Janarthanan Rajendran, Sarath ChandarICLR 2024 · 42 citations
- Parameterized Decision-Making with Multi-Modality Perception for Autonomous DrivingYuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng et al.ICDE 2024 · 26 citations
Builds on10
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
Related papers
- Investigating Memory in RL with POPGym ArcadeZekang Wang, Zhe He, Borong Zhang, Edan Toledo et al.ICML 2026
- Memory Gym: Partially Observable Challenges to Memory-Based AgentsMarco Pleines, Matthias Pallasch, Frank Zimmer, Mike PreussICLR 2023
- Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning EnvironmentsRyan Sullivan, J. K. Terry, Benjamin Black, John P. DickersonICML 2022 · 11 citations
- Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement LearningHung Le, Dung Nguyen, Kien Do, Sunil Gupta et al.ICLR 2025
- Evaluating Long-Term Memory in 3D MazesJurgis Pasukonis, Timothy P. Lillicrap, Danijar HafnerICLR 2023
