AIPO: Adaptive Information Guided Token-Level Reinforcement Learning for Large Language Model Reasoning
Bin Chen, Hongfei Ye, Huiyang Wang, Wenxi Liu, Yu Zhang, Furui Liu
摘要
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning capability of Large Language Models (LLMs). Current RLVR trains LLMs on all generated tokens, rather than exploring which tokens actually contribute to reasoning. We propose AIPO (Adaptive-Information Policy Optimization), which focuses updates on those decisive tokens discovered on the fly. AIPO estimates each hidden state's mutual information to score tokens. Policy gradients are then computed only on these critical tokens, using an advantage that verifiable correctness. To improve the efficiency of mutual-information estimation, AIPO adopts a Random-Fourier approximation of the Hilbert-Schmidt Independence Criterion. Across five math and science benchmarks, AIPO yields up to +20% accuracy over strong RLVR baselines while updating merely 10% of tokens, demonstrating superior efficiency and effectiveness. Our findings highlight the importance of information-driven token selection for efficient and effective reinforcement learning of LLM reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionZhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu 等NeurIPS 2024 · 被引用 125 次
相关 Paper
- IAPO: Information-Aware Policy Optimization for Token-Efficient ReasoningYinhan He, Yaochen Zhu, Mingjia Shi, Wendy Zheng 等ICML 2026 · 被引用 2 次
- Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM ReasoningYiliu Sun, Zicheng Zhao, Yang Wei, Yanfang Zhang 等AAAI 2026 · 被引用 1 次
- TGPO: Efficient Policy Optimization through Sequence Anchor and Information GatingHang Ding, Dongqi Liu, Qiming Feng, Jian Li 等ICML 2026
- Experience Augmented Policy Optimization for LLM ReasoningJinda Lu, Kexin Huang, Junkang Wu, Shuo Yang 等ICML 2026 · 被引用 2 次
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsHaoran He, Yuxiao Ye, Qingpeng Cai, Chen Hu 等ICLR 2026 · 被引用 9 次
