AMRL: Aggregated Memory For Reinforcement Learning
Jacob Beck, Kamil Ciosek, Sam Devlin, Sebastian Tschiatschek, Cheng Zhang, Katja Hofmann
摘要
In many partially observable scenarios, Reinforcement Learning (RL) agents must rely on long-term memory in order to learn an optimal policy. We demonstrate that using techniques from natural language processing and supervised learning fails at RL tasks due to stochasticity from the environment and from exploration. Utilizing our insights on the limitations of traditional memory methods in RL, we propose AMRL, a class of models that can learn better policies with greater sample efficiency and are resilient to noisy inputs. Specifically, our models use a standard memory module to summarize short-term context, and then aggregate all prior states from the standard model without respect to order. We show that this provides advantages both in terms of gradient decay and signal-to-noise ratio over time. Evaluating in Minecraft and maze environments that test long-term memory, we find that our model improves average return by 19% over a baseline that has the same number of parameters and by 9% over a stronger baseline that has far more parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 被引用 77 次
- Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingYaodong Yang, Guangyong Chen, Weixun Wang, Xiaotian Hao 等NeurIPS 2022 · 被引用 24 次
- Recurrent Hypernetworks are Surprisingly Strong in Meta-RLJacob Beck, Risto Vuorio, Zheng Xiong, Shimon WhitesonNeurIPS 2023 · 被引用 22 次
- Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomesArtyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail BurtsevNeurIPS 2022 · 被引用 13 次
相关 Paper
- Semantic HELM: A Human-Readable Memory for Reinforcement LearningFabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp HochreiterNeurIPS 2023 · 被引用 21 次
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal 等NeurIPS 2025 · 被引用 9 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
- Reinforcement Learning with Fast and Forgetful MemorySteven D. Morad, Ryan Kortvelesy, Stephan Liwicki, Amanda ProrokNeurIPS 2023 · 被引用 10 次
- Meta-RL Induces Exploration in Language AgentsYulun Jiang, Liangze Jiang, Damien Teney, Michael Moor 等ICLR 2026 · 被引用 20 次
