Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Cheng Xin, Shuo He, Lang Feng, Haiyang Xu, Ming Yan, Lei Feng, Bo An
摘要
Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agentic tasks. However, their credit assignment relies heavily on coarse-grained trajectory-level attribution according to final outcomes, making it difficult to capture the contribution of individual steps, such as valuable steps obscured within failed trajectories. To uncover latent information and enable more faithful step-level credit assignment, we propose Graph-based Group Policy Optimization (GraphGPO), which first aggregates all rollout trajectories into a unified state-transition graph and then estimates the distance from each state to the task goal using the global information encoded in the graph. Finally, GraphGPO assigns credit to each edge by estimating a graph-based advantage, based on how much the transition reduces the distance to the task goal. In this way, GraphGPO significantly improves training efficiency and achieves state-of-the-art performance across a range of challenging benchmarks. Code is available at https://github.com/langfengQ/ver l-agent/tree/master/recipe/GraphGPO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic TasksShuo He, Lang Feng, Qi Wei, Xin Cheng 等ICLR 2026 · 被引用 36 次
- Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationJunzhe Wang, Zhiheng Xi, Yajie Yang, Hao Luo 等ACL 2026 · 被引用 4 次
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMsYujie Zhao, Lanxiang Hu, Yang Wang, Minmin Hou 等ICLR 2026 · 被引用 26 次
