I²B-LPO: Latent Policy Optimization via Iterative Information Bottleneck
Huilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming Li, Chuyang Zhao, Jihai Zhang, Mengchang Wang, Yang Cao, Yu Kang
Abstract
Despite recent advances in Reinforcement learning with verifiable rewards (RLVR) for large language model (LLM) reasoning, most methods suffer from exploration collapse, as the semantic homogeneity of random rollouts traps models in narrow, over-optimized behaviors. Existing methods leverage policy entropy to encourage exploration, but face inherent limitations: global entropy regularization is susceptible to reward hacking, inducing meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To this end, we propose Latent Policy Optimization via Iterative Information Bottleneck (I 2 B-LPO), which shifts from statistical perturbation of token distributions to topological branching of reasoning trajectories. I 2 B-LPO triggers latent branching at high-entropy states to diversify reasoning trajectories and applies the Information Bottleneck as a trajectory filter and self-reward to ensure concise and informative exploration. Empirical results on four mathematical benchmarks demonstrate that I 2 B-LPO achieves stateof-the-art performance, with margins of up to 5.3% in accuracy and 7.4% in diversity metrics. Code is available at https://github. com/denghuilin-cyber/IIB-LPO . Flatten across the whole vocabulary H z1 H … Critical position Quantile borderline Tokens Tokens (a) Entropy Regularization Method (b) Token-selective Method R: Basically, to actually find the minimum value, well, in the context of optimzation, first, we should differentiate … R: To find the minimum value of the function, let us differentiate the function f(x) with respect to x. (c) I² B-LPO(Ours) … Let us diff... the … Let us diff... the … first we should … … first we should … z2 zi zi zi zi zi zi zi ... Tt Tt+1 ... Tt+1 Tt Tt+n R 1 :we can differentiate the function f(x) with respect to x: R 1 :we can differentiate the function f(x) with respect to x: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R1 R2 To find the minimum value, first, Sharpen on critical position V'
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e357890-c173-4231-b29d-867dfc3bae72Builds on9
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing ReasoningZhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu et al.ICLR 2026 · 271 citations
Related papers
- Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable RewardsShangyu Xing, Siyuan Wang, Chenyuan Yang, Xin-Yu Dai et al.ICLR 2026 · 14 citations
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-ForgetLiang Chen, Xueting Han, Qizhou Wang, Bo Han et al.ICLR 2026 · 16 citations
- ReLaX: Reasoning with Latent Exploration for Large Reasoning ModelsShimin Zhang, Xianwei Chen, Yufan Shen, Ziyuan Ye et al.CVPR 2026 · 4 citations
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy OptimizationHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu et al.ICML 2026
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
