Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, Yunpu Ma
Abstract
Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of NLP tasks, but they remain fundamentally stateless, constrained by limited context windows that hinder long-horizon reasoning. Recent efforts to address this limitation often augment LLMs with an external memory bank, yet most existing pipelines are static and heuristic-driven, lacking a learned mechanism for deciding what to store, update, or retrieve. We present Memory-R1, a reinforcement learning (RL) framework that equips LLMs with the ability to actively manage and utilize external memory through two specialized agents: a Memory Manager that learns structured operations, including ADD, UPDATE, DELETE, and NOOP; and an Answer Agent that pre-selects and reasons over relevant entries. Both agents are fine-tuned with outcome-driven RL (PPO and GRPO), enabling adaptive memory management with minimal supervision. With only 152 training QA pairs, Memory-R1 outperforms strong baselines and generalizes across diverse question types, three benchmarks (LoCoMo, MSC, LongMemEval), and multiple model scales (3B-14B).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a98670e-231a-4a2e-8391-4f739ca00f1cCited by top-tier papers30
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM AgentsYaorui Shi, Yuxin Chen, Siyuan Wang, Sihang Li et al.ICLR 2026 · 35 citations
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng et al.ACL 2026 · 17 citations
- IterResearch: Rethinking Long-Horizon Agents with Interaction ScalingGuoxin Chen, Zile Qiao, Xuanzhong Chen, Donglei Yu et al.ICLR 2026 · 17 citations
- Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningYunghwei Lai, Kaiming Liu, Ziyue Wang, Weizhi Ma et al.ICLR 2026 · 10 citations
- Prompt-MII: Meta-Learning Instruction Induction for LLMsEmily Xiao, Yixiao Zeng, Ada Chen, Chin-Jou Li et al.ICLR 2026 · 9 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
Related papers
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan et al.ACL 2026 · 40 citations
- In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue AgentsZhen Tan, Jun Yan, I-Hung Hsu, Rujun Han et al.ACL 2025
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim et al.ICLR 2026 · 223 citations
- Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory ManagementWeitao Ma, Xiaocheng Feng, Lei Huang, Xiachong Feng et al.ACL 2026 · 2 citations
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li et al.ICLR 2026 · 18 citations
