No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning
Zhicong Li, Lingjie Jiang, Yulan Hu, Xingchen Zeng, Yixia Li, Xiangwen Zhang, Guanhua Chen, Zheng Pan, Xin Li, Yong Liu
Abstract
Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with natural-language feedback. However, current methods often rely on static or offline critic models, which fail to adapt as the policy evolves. In on-policy RL, the agent's error patterns shift over time, causing stationary critics to become stale and providing feedback of diminishing utility. To address this, we introduce ECHO (Evolving Critic for Hindsight-Guided Optimization), a framework that jointly optimizes the policy and critic through a synchronized co-evolutionary loop. ECHO utilizes a cascaded rollout mechanism where the critic generates multiple diagnoses for an initial trajectory, followed by policy refinement to enable group-structured advantage estimation. We address the challenge of learning plateaus via a saturation-aware gain shaping objective, which rewards the critic for inducing incremental improvements in high-performing trajectories. By employing dual-track GRPO updates, ECHO ensures the critic's feedback stays synchronized with the evolving policy. Experimental results show that ECHO yields more stable training and higher long-horizon task success across open-world environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14fec7fe-a871-4f93-9674-240f5d0f890aCited by top-tier papers1
Ask how each one uses itBuilds on10
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang et al.NeurIPS 2025 · 310 citations
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM ReasoningXichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan et al.ICLR 2026 · 52 citations
Related papers
- Advancing LLM Reasoning with Natural Language and Numerical FeedbackXiaoying Zhang, Yipeng Zhang, Hao Sun, Kaituo Feng et al.ICML 2026 · 79 citations
- Self-evolving LLM agents with in-distribution OptimizationYudi Zhang, Meng Fang, Zhenfang Chen, Mykola PechenizkiyICML 2026 · 1 citation
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan et al.NeurIPS 2025 · 21 citations
- Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context ReasoningXin Guan, Zijian Li, Shen Huang, Pengjun Xie et al.ACL 2026 · 6 citations
- Critique-RL: Training Language Models For Critiquing Through Two-Stage Reinforcement LearningZhiheng Xi, Jixuan Huang, Xin Guo, Boyang Hong et al.ICLR 2026 · 4 citations
