Recurrent Action Transformer with Memory
Egor Cherepanov, Aleksei Staroverov, Alexey Kovalev, Aleksandr Panov
Abstract
Transformers have become increasingly popular in offline reinforcement learning (RL) due to their ability to treat agent trajectories as sequences, reframing policy learning as a sequence modeling task. However, in partially observable environments (POMDPs), effective decision-making depends on retaining information about past events -something that standard transformers struggle with due to the quadratic complexity of self-attention, which limits their context length. One solution to this problem is to extend transformers with memory mechanisms. We propose the Recurrent Action Transformer with Memory (RATE), a novel transformer-based architecture for offline RL that incorporates a recurrent memory mechanism designed to regulate information retention. We evaluate RATE across a diverse set of environments: memory-intensive tasks (ViZDoom-Two-Colors, T-Maze, Memory Maze, Minigrid-Memory, and POP-Gym), as well as standard Atari and MuJoCo benchmarks. Our comprehensive experiments demonstrate that RATE significantly improves performance in memory-dependent settings while remaining competitive on standard tasks across a broad range of baselines. These findings underscore the pivotal role of integrated memory mechanisms in offline RL and establish RATE as a unified, highcapacity architecture for effective decision-making over extended horizons. Code: https://sites.google.com/view/rate-model/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31d491f1-a0a1-48df-bc8f-385dd5549344Cited by top-tier papers5
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningEgor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. PanovICLR 2026 · 43 citations
- Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and EvaluationEgor Cherepanov, Nikita Kachaev, Artem Zholus, Alexey K. Kovalev et al.ICLR 2026 · 4 citations
- Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsShuai Zhen, Yanhua Yu, Ruopei Guo, Nan Cheng et al.ACL 2026 · 2 citations
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL ProblemsEgor Cherepanov, Alexey Kovalev, Aleksandr PanovICLR 2026 · 1 citation
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal DemonstrationsAnian Ruoss, Fabio Pardo, Harris Chan, Bonnie Li et al.ICML 2025
Builds on24
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal et al.NeurIPS 2025 · 9 citations
- Decision S4: Efficient Sequence-Based RL via State Spaces LayersShmuel Bar-David, Itamar Zimerman, Eliya Nachmani, Lior WolfICLR 2023 · 3 citations
- Rethinking Transformers in Solving POMDPsChenhao Lu, Ruizhe Shi, Yuyao Liu, Kaizhe Hu et al.ICML 2024 · 10 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- Transformers are Meta-Reinforcement LearnersLuckeciano C. MeloICML 2022 · 66 citations
