No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning
Zhicong Li, Lingjie Jiang, Yulan Hu, Xingchen Zeng, Yixia Li, Xiangwen Zhang, Guanhua Chen, Zheng Pan, Xin Li, Yong Liu
摘要
Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with natural-language feedback. However, current methods often rely on static or offline critic models, which fail to adapt as the policy evolves. In on-policy RL, the agent's error patterns shift over time, causing stationary critics to become stale and providing feedback of diminishing utility. To address this, we introduce ECHO (Evolving Critic for Hindsight-Guided Optimization), a framework that jointly optimizes the policy and critic through a synchronized co-evolutionary loop. ECHO utilizes a cascaded rollout mechanism where the critic generates multiple diagnoses for an initial trajectory, followed by policy refinement to enable group-structured advantage estimation. We address the challenge of learning plateaus via a saturation-aware gain shaping objective, which rewards the critic for inducing incremental improvements in high-performing trajectories. By employing dual-track GRPO updates, ECHO ensures the critic's feedback stays synchronized with the evolving policy. Experimental results show that ECHO yields more stable training and higher long-horizon task success across open-world environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM ReasoningXichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan 等ICLR 2026 · 被引用 52 次
相关 Paper
- Advancing LLM Reasoning with Natural Language and Numerical FeedbackXiaoying Zhang, Yipeng Zhang, Hao Sun, Kaituo Feng 等ICML 2026 · 被引用 79 次
- Self-evolving LLM agents with in-distribution OptimizationYudi Zhang, Meng Fang, Zhenfang Chen, Mykola PechenizkiyICML 2026 · 被引用 1 次
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan 等NeurIPS 2025 · 被引用 21 次
- Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context ReasoningXin Guan, Zijian Li, Shen Huang, Pengjun Xie 等ACL 2026 · 被引用 6 次
- Critique-RL: Training Language Models For Critiquing Through Two-Stage Reinforcement LearningZhiheng Xi, Jixuan Huang, Xin Guo, Boyang Hong 等ICLR 2026 · 被引用 4 次
