On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
Deyu Zou, Yongqiang Chen, Fan Feng, Mufei Li, Pan Li, Yu Gong, James Cheng
摘要
Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons. However, in active reasoning where agents must elicit new observations through interaction with the environment to solve the task, we find that outcome-based RL can induce a systematic failure mode which we call information self-locking (SeL): agents fail both to elicit informative feedback and to internalize obtained evidence. To understand the issue, we trace agentic behaviors into two coupled capabilities: Action Selection (AS), which determines observation streams, and Belief Tracking (BT), which updates the agent's internal task understanding. Theoretical and empirical analyses reveal a bidirectional bottleneck that leads to SeL: weak BT obscures the credit of informative actions, while weak AS deprives BT of useful evidence. This coupling weakens the learning signal for both capabilities and leads to SeL. To mitigate this issue, we propose AREW, a simple yet effective Advantage-Reweighting method that uses easy-to-obtain directional critiques to reallocate credit within trajectories. Extensive experiments across 9 agentic tasks of varying complexity show that AREW significantly mitigates SeL, yielding up to 60-point gains in final performance. Code is available at https://github.com/unimpor/T3 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CausalGame: Benchmarking Causal Thinking of LLM Agents in GamesZhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu 等ICML 2026 · 被引用 2 次
- Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement LearningZizhe Chen, Jiqian Dong, Yizhou Tian, Garry YANG 等ICML 2026 · 被引用 1 次
- Concept Concentration for Faithful Representation InterventionHongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu 等ICML 2026
它引用的顶会 Paper5
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsZhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao 等NeurIPS 2024 · 被引用 43 次
- Enhancing Personalized Multi-Turn Dialogue with Curiosity RewardYanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani 等NeurIPS 2025 · 被引用 32 次
- End-to-End Learning of Flowchart Grounded Task-Oriented DialogsDinesh Raghu, Shantanu Agarwal, Sachindra Joshi, MausamEMNLP 2021 · 被引用 12 次
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang 等ICML 2026 · 被引用 3 次
相关 Paper
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM AgentsDeyu Zou, Yongqiang Chen, Jianxiang Wang, Garry Yang 等ICLR 2026 · 被引用 3 次
- SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving AgentsXinshun Feng, Xinhao Song, Lijun Li, Gongshen Liu 等ACL 2026 · 被引用 2 次
- Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM ReasoningQiao Liang, Yuke Zhu, Chao Ge, Lei Yang 等ACL 2026 · 被引用 4 次
- SOAR: Supervision from Observation for Agentic Reinforcement LearningMeng Li, Lei Li, Xiting Wang, Yi Yuan 等ACL 2026 · 被引用 1 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
