Lune

ICML2026顶会

Understanding Reasoning Collapse in LLM Agent Reinforcement Learning

Zihan (Zenus) Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu

出版方
2026年份

摘要

In closed-loop multi-turn agent reinforcement learning, LLM agents exhibit reasoning collapse, where reasoning shift toward generic templates, weakly coupled to the inputs. We firstly identify that such collapse is easy to miss with entropy or surface diversity metrics since reasoning text still varies but becomes input-agnostic. We then propose an information-theoretic decomposition of reasoning variable ZZ's variation into conditional entropy H(Z∣X)H(Z \mid X) (randomness under same input) and mutual information (MI) I(X;Z)I(X; Z) (input dependence). Template collapse occurs when H(Z∣X)H(Z \mid X) stays high while I(X;Z)I(X; Z) drops, yielding diverse-looking but generic reasoning. To make I(X;Z)I(X; Z) a reproducible and sanity-checkable diagnostic, we further introduce an MI-style retrieval protocol treating each reasoning trace ZZ as a query to retrieve its source XX from a minibatch; accuracy degrades toward chance under collapse. We thus provide a signal-to-noise ratio explanation for why I(X;Z)I(X; Z) drops: when within-input reward variance Var(R∣X)\mathrm{Var}(R \mid X) is low, task gradients weaken and input-agnostic regularizers (KL, entropy) dominate, flattening cross-input differences. Finally, we propose reward-variance-aware filtering to prioritize high-signal updates. Across multi-turn environments, model scales, and modalities (including VLMs), this improves input dependence, stability, and performance while remaining competitive with state-of-the-art stabilization baselines.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖