MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
Zhiyu Shen, Ziming Wu, Fuming Lai, Shaobing Lian, Yanghui Rao
摘要
Maintaining consistency in long-term dialogues remains a fundamental challenge for LLMs, as standard retrieval mechanisms often fail to capture the temporal evolution of historical states. While memory-augmented frameworks offer a structured alternative, current systems rely on static prompting of closed-source models or suffer from ineffective training paradigms with sparse rewards. We introduce MemBuilder, a reinforcement learning framework that trains models to orchestrate multi-dimensional memory construction with attributed dense rewards. MemBuilder addresses two key challenges: (1) Sparse Trajectory-Level Rewards: we employ synthetic session-level question generation to provide dense intermediate rewards across extended trajectories; and (2) Multi-Dimensional Memory Attribution: we introduce contribution-aware gradient weighting that scales policy updates based on each component's downstream impact. Experimental results show that MemBuilder enables a 4Bparameter model to outperform state-of-the-art closed-source baselines, exhibiting strong generalization across long-term dialogue benchmarks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- Augmenting Language Models with Long-Term MemoryWeizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu 等NeurIPS 2023 · 被引用 256 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim 等ICLR 2026 · 被引用 223 次
相关 Paper
- From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational AgentsDerong Xu, Yi Wen, Pengyue Jia, Yingyi Zhang 等ICLR 2026 · 被引用 28 次
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session AgentsYiming Du, Baojun Wang, Yifan Xiang, Zhaowei Wang 等ICLR 2026 · 被引用 11 次
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM AgentsYaorui Shi, Yuxin Chen, Siyuan Wang, Sihang Li 等ICLR 2026 · 被引用 35 次
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie 等ACL 2026 · 被引用 140 次
- M+: Extending MemoryLLM with Scalable Long-Term MemoryYu Wang, Dmitry Krotov, Yuanzhe Hu, Yifan Gao 等ICML 2025
