Large Language Models Are Semi-Parametric Reinforcement Learning Agents
Danyang Zhang, Lu Chen, Situo Zhang, Hongshen Xu, Zihan Zhao, Kai Yu
Abstract
Inspired by the insights in cognitive science with respect to human memory and reasoning mechanism, a novel evolvable LLM-based (Large Language Model) agent framework is proposed as REMEMBERER. By equipping the LLM with a long-term experience memory, REMEMBERER is capable of exploiting the experiences from the past episodes even for different task goals, which excels an LLM-based agent with fixed exemplars or equipped with a transient working memory. We further introduce Reinforcement Learning with Experience Memory (RLEM) to update the memory. Thus, the whole system can learn from the experiences of both success and failure, and evolve its capability without fine-tuning the parameters of the LLM. In this way, the proposed REMEMBERER constitutes a semi-parametric RL agent. Extensive experiments are conducted on two RL task sets to evaluate the proposed framework. The average results with different initialization and training sets exceed the prior SOTA by 4% and 2% for the success rate on two task sets and demonstrate the superiority and robustness of REMEMBERER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 069cfa34-939c-4069-a0d0-95019b8ffbf3Cited by top-tier papers14
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 · 48 citations
- Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement LearningYun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao et al.AAAI 2025 · 29 citations
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li et al.ICLR 2026 · 18 citations
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy OptimizationChengyu Huang, Zhengxin Zhang, Claire CardieAAAI 2026 · 14 citations
- Tunable LLM-based Proactive Recommendation AgentMingze Wang, Chongming Gao, Wenjie Wang, Yangyang Li et al.ACL 2025 · 4 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 · 102 citations
- UMEM: Unified Memory Extraction and Management Framework for Generalizable MemoryYongshi Ye, Hui Jiang, Feihu Jiang, Tian Lan et al.ICML 2026 · 4 citations
- Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience DistillationZihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang et al.ACL 2026 · 4 citations
- MemEvolve: Meta-Evolution of Agent Memory SystemsGuibin Zhang, Haotian Ren, Chong Zhan, Junhao Wang et al.ICML 2026 · 69 citations
- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language ModelMengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu et al.ACL 2025
