Large Language Models Are Semi-Parametric Reinforcement Learning Agents
Danyang Zhang, Lu Chen, Situo Zhang, Hongshen Xu, Zihan Zhao, Kai Yu
摘要
Inspired by the insights in cognitive science with respect to human memory and reasoning mechanism, a novel evolvable LLM-based (Large Language Model) agent framework is proposed as REMEMBERER. By equipping the LLM with a long-term experience memory, REMEMBERER is capable of exploiting the experiences from the past episodes even for different task goals, which excels an LLM-based agent with fixed exemplars or equipped with a transient working memory. We further introduce Reinforcement Learning with Experience Memory (RLEM) to update the memory. Thus, the whole system can learn from the experiences of both success and failure, and evolve its capability without fine-tuning the parameters of the LLM. In this way, the proposed REMEMBERER constitutes a semi-parametric RL agent. Extensive experiments are conducted on two RL task sets to evaluate the proposed framework. The average results with different initialization and training sets exceed the prior SOTA by 4% and 2% for the success rate on two task sets and demonstrate the superiority and robustness of REMEMBERER.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 被引用 48 次
- Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement LearningYun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao 等AAAI 2025 · 被引用 29 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy OptimizationChengyu Huang, Zhengxin Zhang, Claire CardieAAAI 2026 · 被引用 14 次
- Tunable LLM-based Proactive Recommendation AgentMingze Wang, Chongming Gao, Wenjie Wang, Yangyang Li 等ACL 2025 · 被引用 4 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 · 被引用 102 次
- UMEM: Unified Memory Extraction and Management Framework for Generalizable MemoryYongshi Ye, Hui Jiang, Feihu Jiang, Tian Lan 等ICML 2026 · 被引用 4 次
- Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience DistillationZihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang 等ACL 2026 · 被引用 4 次
- MemEvolve: Meta-Evolution of Agent Memory SystemsGuibin Zhang, Haotian Ren, Chong Zhan, Junhao Wang 等ICML 2026 · 被引用 69 次
- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language ModelMengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu 等ACL 2025
