Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural Memory
Sijia Li, Yuchen Huang, Zifan LIU, Zijian LI, Jingjing Fu, Lei Song, Jiang Bian, Jun Zhang, Rui Wang
摘要
As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience is intuitively appealing, existing approaches remain limited: full trajectories are often too context-specific to transfer, while tool-level reuse ignores the context and environment. In this paper, we introduce a hybrid episodic–procedural memory strategy (H-EPM) that enables experience-evolution of multi-turn tool-use policies, by adaptively reusing partially overlapping successful experiences in both inference and training. Inspired by human episodic–procedural integration, we build a tool graph from accumulated trajectories, where recurring tool-to-tool dependencies capture procedural routines and each edge is augmented with a compact episodic summaries of relevant context. At inference, the agent dynamically balances episodic recall for contextual reasoning and procedural execution for routine steps. Beyond inference, H-EPM introduces a memory-guided reinforcement learning paradigm that directly addresses a core challenge in multi-turn agent RL: ineffective exploration over long trajectories. By biasing exploration toward historically successful tool transitions, H-EPM learns a stronger policy that generalizes during inference without relying on domain-specific experience collection. Experiments show that H-EPM consistently delivers substantial inference-time gains over strong baselines across multi-turn tool-use benchmarks, reaching up to 50%+. It also boosts RL policy performance, achieving up to 40%+ improvement on out-of-distribution tasks. Our code is available at https://github.com/LISijia-dev/H-EPM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemorySiru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen 等ICLR 2026 · 被引用 244 次
- A Human-Inspired Reading Agent with Gist Memory of Very Long ContextsKuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John F. Canny 等ICML 2024 · 被引用 106 次
- Can Graph Learning Improve Planning in LLM-based Agents?Xixi Wu, Yifei Shen, Caihua Shan, Kaitao Song 等NeurIPS 2024 · 被引用 67 次
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following BehaviorZidi Xiong, Yuping Lin, Wenya Xie, Pengfei He 等ACL 2026 · 被引用 64 次
相关 Paper
- REMem: Reasoning with Episodic Memory in Language AgentYiheng Shu, Padmaja Jonnalagedda, Xiang Gao, Bernal Jimenez Gutierrez 等ICLR 2026 · 被引用 20 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
- Model-Based Episodic Memory Induces Dynamic Hybrid ControlsHung Le, Thommen George Karimpanal, Majid Abdolshah, Truyen Tran 等NeurIPS 2021 · 被引用 25 次
- From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM AgentsRong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai 等ICML 2026
- ProcMEM: Learning Reusable Procedural Memory from Experience via Non-Parametric PPO for LLM AgentsQIRUI MI, Zhijian Ma, Mengyue Yang, Yisen Wang 等ICML 2026
