IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
Yinhan He, Yaochen Zhu, Mingjia Shi, Wendy Zheng, Lin Su, Xiaoqing Wang, Qi Guo, Jundong Li
Abstract
Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We revisit token-efficient post-training and argue that existing sequence-level reward-shaping methods offer limited control over how reasoning effort is allocated across tokens. To bridge the gap, we propose IAPO, an information-theoretic post-training framework that assigns token-wise advantages based on each token’s conditional mutual information (MI) with the final answer. This yields an explicit, principled mechanism for identifying informative reasoning steps and suppressing low-utility exploration. We provide a theoretical analysis showing that our IAPO can induce monotonic reductions in reasoning verbosity without harming correctness. Empirically, IAPO consistently improves reasoning accuracy while reducing reasoning length by up to 36%, outperforming existing token-efficient RL methods across various reasoning datasets. Our results demonstrate that information-aware advantage shaping is a powerful and general direction for token-efficient post-training. The code is available at https://github.com/YinhanHe123/IAPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0220c0ab-b78b-41f3-b44b-99684373971eCited by top-tier papers1
Ask how each one uses itBuilds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
Related papers
- AIPO: Adaptive Information Guided Token-Level Reinforcement Learning for Large Language Model ReasoningBin Chen, Hongfei Ye, Huiyang Wang, Wenxi Liu et al.ACL 2026
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou et al.ACL 2026 · 5 citations
- Step-GRPO: Internalizing Dynamic Early Exit for Efficient ReasoningBenteng Chen, Weida Wang, Shufei Zhang, Mingbao Lin et al.ACL 2026
- SSVPO: Effective Step-Level Credit Assignment for RL Training of Language ModelsYugu Li, Zehong Cao, Jianglin Qiao, Siyi HuICLR 2026
- Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning ModelsWei Wu, Liyi Chen, Congxi Xiao, Tianfu Wang et al.ACL 2026 · 1 citation
