On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
Deyu Zou, Yongqiang Chen, Fan Feng, Mufei Li, Pan Li, Yu Gong, James Cheng
Abstract
Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons. However, in active reasoning where agents must elicit new observations through interaction with the environment to solve the task, we find that outcome-based RL can induce a systematic failure mode which we call information self-locking (SeL): agents fail both to elicit informative feedback and to internalize obtained evidence. To understand the issue, we trace agentic behaviors into two coupled capabilities: Action Selection (AS), which determines observation streams, and Belief Tracking (BT), which updates the agent's internal task understanding. Theoretical and empirical analyses reveal a bidirectional bottleneck that leads to SeL: weak BT obscures the credit of informative actions, while weak AS deprives BT of useful evidence. This coupling weakens the learning signal for both capabilities and leads to SeL. To mitigate this issue, we propose AREW, a simple yet effective Advantage-Reweighting method that uses easy-to-obtain directional critiques to reallocate credit within trajectories. Extensive experiments across 9 agentic tasks of varying complexity show that AREW significantly mitigates SeL, yielding up to 60-point gains in final performance. Code is available at https://github.com/unimpor/T3 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f91f372-ea44-4467-bd4f-893c8de9d1fcCited by top-tier papers3
- CausalGame: Benchmarking Causal Thinking of LLM Agents in GamesZhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu et al.ICML 2026 · 2 citations
- Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement LearningZizhe Chen, Jiqian Dong, Yizhou Tian, Garry YANG et al.ICML 2026 · 1 citation
- Concept Concentration for Faithful Representation InterventionHongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu et al.ICML 2026
Builds on5
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsZhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao et al.NeurIPS 2024 · 43 citations
- Enhancing Personalized Multi-Turn Dialogue with Curiosity RewardYanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani et al.NeurIPS 2025 · 32 citations
- End-to-End Learning of Flowchart Grounded Task-Oriented DialogsDinesh Raghu, Shantanu Agarwal, Sachindra Joshi, MausamEMNLP 2021 · 12 citations
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang et al.ICML 2026 · 3 citations
Related papers
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning of LLM AgentsDeyu Zou, Yongqiang Chen, Jianxiang Wang, Garry Yang et al.ICLR 2026 · 3 citations
- SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving AgentsXinshun Feng, Xinhao Song, Lijun Li, Gongshen Liu et al.ACL 2026 · 2 citations
- Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM ReasoningQiao Liang, Yuke Zhu, Chao Ge, Lei Yang et al.ACL 2026 · 4 citations
- SOAR: Supervision from Observation for Agentic Reinforcement LearningMeng Li, Lei Li, Xiting Wang, Yi Yuan et al.ACL 2026 · 1 citation
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
