Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural Memory
Sijia Li, Yuchen Huang, Zifan LIU, Zijian LI, Jingjing Fu, Lei Song, Jiang Bian, Jun Zhang, Rui Wang
Abstract
As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing approaches remain limited: full trajectories are often too context-specific to transfer, while tool-level reuse ignores the context and environment. In this paper, we introduce a hybrid episodic–procedural memory strategy (H-EPM) that enables experience-evolution of multi-turn tool-use policies, by adaptively reusing partially overlapping successful experiences in both inference and training. Inspired by human episodic–procedural integration, we build a tool graph from accumulated trajectories, where recurring tool-to-tool dependencies capture procedural routines and each edge is augmented with a compact episodic summaries of relevant context. At inference, the agent dynamically balances episodic recall for contextual reasoning and procedural execution for routine steps. Beyond inference, H-EPM introduces a memory-guided reinforcement learning paradigm that directly addresses a core challenge in multi-turn agent RL: ineffective exploration over long trajectories. By biasing exploration toward historically successful tool transitions, H-EPM learns a stronger policy that generalizes during inference without relying on domain-specific experience collection. Experiments show that H-EPM consistently delivers substantial inference-time gains over strong baselines across multi-turn tool-use benchmarks, reaching up to 50%+. It also boosts RL policy performance, achieving up to 40%+ improvement on out-of-distribution tasks. Our code is available at https://github.com/LISijia-dev/H-EPM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c23955e-f7b3-4382-b98e-fa7e1f7c413fBuilds on10
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemorySiru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen et al.ICLR 2026 · 244 citations
- A Human-Inspired Reading Agent with Gist Memory of Very Long ContextsKuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John F. Canny et al.ICML 2024 · 106 citations
- Can Graph Learning Improve Planning in LLM-based Agents?Xixi Wu, Yifei Shen, Caihua Shan, Kaitao Song et al.NeurIPS 2024 · 67 citations
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following BehaviorZidi Xiong, Yuping Lin, Wenya Xie, Pengfei He et al.ACL 2026 · 64 citations
Related papers
- REMem: Reasoning with Episodic Memory in Language AgentYiheng Shu, Padmaja Jonnalagedda, Xiang Gao, Bernal Jimenez Gutierrez et al.ICLR 2026 · 20 citations
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li et al.ICLR 2026 · 18 citations
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran et al.NeurIPS 2021 · 25 citations
- From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM AgentsRong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai et al.ICML 2026
- ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM AgentsQIRUI MI, Zhijian Ma, Mengyue Yang, Yisen Wang et al.ICML 2026
