Self-evolving LLM agents with in-distribution Optimization
Yudi Zhang, Meng Fang, Zhenfang Chen, Mykola Pechenizkiy
Abstract
Large Language Models (LLMs) have recently emerged as powerful controllers for interactive agents in complex environments, yet training them to perform reliable long-horizon decision making remains a fundamental challenge. A key difficulty lies in credit assignment: agents often receive delayed rewards only at the end of episodes. In this paper, we propose Q-Evolve, a self-evolving framework for LLM agents that unifies automatic process-reward labeling and policy learning within a principled in-distribution reinforcement learning paradigm. In each evolving iteration, our method learns an in-distribution critic from a hybrid off-policy dataset that combines expert demonstrations with agent-generated trajectories, stabilizing Bellman backups in sparse-reward settings via a weighted Implicit Q-Learning objective. The learned value function is then used to derive step-wise process rewards through advantage estimation, enabling dense and reliable supervision without environment backtracking or human annotation. Leveraging these signals, we perform behavior-proximal policy optimization that evolves the agent over the data used for process reward labeling, allowing iterative self-improvement without exacerbating distribution shift. We evaluate our method on AlfWorld, WebShop, and ScienceWorld, showing Q-Evolve outperforms strong baselines in sample efficiency, robustness, and overall task performance. Our results demonstrate that stable agent self-evolution is achievable through the co-evolution of process-level supervision and policy, both grounded within a shared in-distribution learning loop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc90df6c-ae40-4636-bdcf-4ee2b8c63598Builds on34
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
Related papers
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang et al.ICLR 2026 · 46 citations
- CoEvolve: Training LLM Agents via Agent-Data Mutual EvolutionShidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu et al.ACL 2026 · 6 citations
- Retrospective In-Context Learning for Temporal Credit Assignment with Large Language ModelsWen-Tse Chen, Jiayu Chen, Fahim Tajwar, Hao Zhu et al.NeurIPS 2025 · 4 citations
- REvolve: Reward Evolution with Large Language Models using Human FeedbackRishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi et al.ICLR 2025
- SHADOW: Dynamic-Aware Credit Assignment Against Long-Horizon TasksYuze Liu, Chaochao Lu, Chao YangAAAI 2026
