SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
Xinshun Feng, Xinhao Song, Lijun Li, Gongshen Liu, Jing Shao
Abstract
Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks. With the paradigm shift toward self-evolving agentic learning, models are increasingly expected to learn from trajectories by synthesizing tools or accumulating explicit experiences. However, prevailing methods typically rely on large-scale LLMs or multi-agent frameworks, which hinder their deployment in resource-constrained environments. The inherent sparsity of outcome-based rewards also poses a substantial challenge, as agents typically receive feedback only upon completion of tasks. To address these limitations, we introduce a Tool-Memory based self-evolving agentic framework SEARL. Unlike approaches that directly utilize interaction experiences, our method constructs a structured experience memory that integrates planning with execution. This provides a novel state abstraction that facilitates generalization across analogous contexts, such as tool reuse. Consequently, agents extract explicit knowledge from historical data while leveraging inter-trajectory correlations to densify reward signals. We evaluate our framework on knowledge reasoning and mathematics tasks, demonstrating its effectiveness in achieving more practical and efficient learning 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e79d3c8-8485-43c0-99f3-05bf834d3887Builds on7
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
- Large Language Models Are Semi-Parametric Reinforcement Learning AgentsDanyang Zhang, Lu Chen, Situo Zhang, Hongshen Xu et al.NeurIPS 2023 · 56 citations
- From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven InteractionsChangle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai et al.ICLR 2025
Related papers
- From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM AgentsRong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai et al.ICML 2026
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
- Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem SolvingXinji Mai, Haotian Xu, Xing W, Weinong Wang et al.NeurIPS 2025 · 7 citations
- UMEM: Unified Memory Extraction and Management Framework for Generalizable MemoryYongshi Ye, Hui Jiang, Feihu Jiang, Tian Lan et al.ICML 2026 · 4 citations
- Evolving AgentsLeonardo RanaldiACL 2026 · 227 citations
