Understanding Reasoning Collapse in LLM Agent Reinforcement Learning
Zihan (Zenus) Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu
摘要
In closed-loop multi-turn agent reinforcement learning, LLM agents exhibit reasoning collapse, where reasoning shift toward generic templates, weakly coupled to the inputs. We firstly identify that such collapse is easy to miss with entropy or surface diversity metrics since reasoning text still varies but becomes input-agnostic. We then propose an information-theoretic decomposition of reasoning variable 's variation into conditional entropy (randomness under same input) and mutual information (MI) (input dependence). Template collapse occurs when stays high while drops, yielding diverse-looking but generic reasoning. To make a reproducible and sanity-checkable diagnostic, we further introduce an MI-style retrieval protocol treating each reasoning trace as a query to retrieve its source from a minibatch; accuracy degrades toward chance under collapse. We thus provide a signal-to-noise ratio explanation for why drops: when within-input reward variance is low, task gradients weaken and input-agnostic regularizers (KL, entropy) dominate, flattening cross-input differences. Finally, we propose reward-variance-aware filtering to prioritize high-signal updates. Across multi-turn environments, model scales, and modalities (including VLMs), this improves input dependence, stability, and performance while remaining competitive with state-of-the-art stabilization baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- ANCHOR: Taming Entropy Dynamics for Stable and Efficient Reasoning of Large Language ModelsCong Qin, Jiaye Lin, Xiaoliang Fu, Yangyi Fang 等KDD 2026
- VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin 等ACL 2026 · 被引用 9 次
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
- The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMsZihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding 等ACL 2026
- A Narrowing Geometry in Contaminated ReasoningJiakuan Xie, Pengfei Cao, Kang Liu, Jun ZhaoICML 2026
