Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
Jingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng, Hui Xiong
摘要
Large language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectiveness and efficiency, often encouraging unnecessarily long reasoning chains and wasting tokens. To address this, we propose Learning to Think (L2T) 3 , an information-theoretic reinforcement fine-tuning framework for LLMs to make the models achieve optimal reasoning with fewer tokens. Specifically, L2T treats each query-response interaction as a hierarchical session of multiple episodes and proposes a universal dense process reward, i.e., quantifies the episode-wise information gain in parameters, requiring no extra annotations or task-specific evaluators. We propose a method to quickly estimate this reward based on PAC-Bayes bounds and the Fisher information matrix. Theoretical analyses show that it significantly reduces computational complexity with high estimation accuracy. By immediately rewarding each episode's contribution and penalizing excessive updates, L2T optimizes the model via reinforcement learning to maximize the use of each episode and achieve effective updates. Empirical results on various reasoning benchmarks and base models demonstrate the advantage of L2T across different tasks, boosting both reasoning effectiveness and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection SuppressionJiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen 等AAAI 2026 · 被引用 21 次
- Rectifying LLM Thought from Lens of OptimizationJunnan Liu, Hongwei Liu, Songyang Zhang, Kai ChenICLR 2026 · 被引用 3 次
- On the Plasticity and Stability for Post-Training Large Language ModelsWenwen Qiang, Ziyin Gu, Jiahuan Zhou, Jie Hu 等ICML 2026 · 被引用 3 次
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsPeizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou 等CVPR 2026 · 被引用 1 次
- A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and SolutionsZhiyin Yu, Yuchen Mou, Juncheng Yan, Junyu Luo 等ACL 2026
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou 等ACL 2026 · 被引用 5 次
- Learning to Reason Efficiently with Discounted Reinforcement LearningAlex Ayoub, Kavosh Asadi, Dale Schuurmans, Csaba Szepesvari 等ICLR 2026 · 被引用 4 次
- Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty QuantificationShaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong WangICML 2026
- Optimizing Test-Time Compute via Meta Reinforcement FinetuningYuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall 等ICML 2025
- Incentivizing LLM Reasoning via Reinforcement Learning with Functional Monte Carlo Tree SearchKongcheng Zhang, QI YAO, Baisheng Lai, Jiaxing Huang 等ICLR 2026
