On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
wenlong deng, Yushu Li, Boying Gong, YI REN, Christos Thrampoulidis, Xiaoxiao Li
摘要
Tool-integrated (TI) reinforcement learning (RL) enables large language models (LLMs) to perform multi-step reasoning by interacting with external tools such as search engines and retrievers. Group Relative Policy Optimization (GRPO), exemplified by the recent Search-R1, offers fast convergence and a value-free formulation that makes it appealing for this setting, yet consistently suffers from training collapse. We identify Lazy Likelihood Displacement (LLD), a systematic reduction or stagnation in the likelihood of both correct and incorrect responses, as the core mechanism driving this failure. LLD emerges early and triggers a self-reinforcing LLD Death Spiral, where declining likelihood leads to low-confidence responses, inflating gradients, and ultimately causing collapse. We empirically characterize this process across models on a Search-R1-style, search-integrated question answering task, revealing a consistent three-phase trajectory: early stagnation, steady decay, and accelerated collapse. To address this, we propose a likelihood-preserving regularization LLDS that activates only when a response action’s likelihood decreases, and regularizes only the tokens responsible. This fine-grained structure mitigates LLD with minimal interference. Our method stabilizes training, prevents gradient explosion, and yields substantial performance improvements across seven benchmarks, including relative improvements of +45.2% on Qwen2.5-3B and +37.1% on Qwen2.5-7B over vanilla GRPO training. Our results establish LLD as a previously overlooked bottleneck in GRPO- based TIRL and provide a practical path toward stable, scalable training of tool-integrated RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ETS: Energy-Guided Test-Time Scaling for Training-Free RL AlignmentXiuyu Li, Jinkai Zhang, Mingyang Yi, Yu Li 等ICML 2026 · 被引用 4 次
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward DecompositionYanyu Chen, Jiyue Jiang, Dianzhi Yu, Zheng Wu 等KDD 2026
它引用的顶会 Paper11
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley 等NeurIPS 2023 · 被引用 515 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningZhenghai Xue, Longtao Zheng, Qian Liu, Yingru Li 等ICLR 2026 · 被引用 152 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
相关 Paper
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationWenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland 等NeurIPS 2025 · 被引用 36 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability ScalingAbhijeet Sinha, Sundari Elango, Dianbo LiuICML 2026 · 被引用 2 次
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han 等ICLR 2026 · 被引用 47 次
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihoodXingyu Lin, Yilin Wen, Du Su, En Wang 等ACL 2026
