I²B-LPO: Latent Policy Optimization via Iterative Information Bottleneck
Huilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming Li, Chuyang Zhao, Jihai Zhang, Mengchang Wang, Yang Cao, Yu Kang
摘要
Despite recent advances in Reinforcement learning with verifiable rewards (RLVR) for large language model (LLM) reasoning, most methods suffer from exploration collapse, as the semantic homogeneity of random rollouts traps models in narrow, over-optimized behaviors. Existing methods leverage policy entropy to encourage exploration, but face inherent limitations: global entropy regularization is susceptible to reward hacking, inducing meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To this end, we propose Latent Policy Optimization via Iterative Information Bottleneck (I 2 B-LPO), which shifts from statistical perturbation of token distributions to topological branching of reasoning trajectories. I 2 B-LPO triggers latent branching at high-entropy states to diversify reasoning trajectories and applies the Information Bottleneck as a trajectory filter and self-reward to ensure concise and informative exploration. Empirical results on four mathematical benchmarks demonstrate that I 2 B-LPO achieves stateof-the-art performance, with margins of up to 5.3% in accuracy and 7.4% in diversity metrics. Code is available at https://github. com/denghuilin-cyber/IIB-LPO . Flatten across the whole vocabulary H z1 H … Critical position Quantile borderline Tokens Tokens (a) Entropy Regularization Method (b) Token-selective Method R: Basically, to actually find the minimum value, well, in the context of optimzation, first, we should differentiate … R: To find the minimum value of the function, let us differentiate the function f(x) with respect to x. (c) I² B-LPO(Ours) … Let us diff... the … Let us diff... the … first we should … … first we should … z2 zi zi zi zi zi zi zi ... Tt Tt+1 ... Tt+1 Tt Tt+n R 1 :we can differentiate the function f(x) with respect to x: R 1 :we can differentiate the function f(x) with respect to x: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R1 R2 To find the minimum value, first, Sharpen on critical position V'
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing ReasoningZhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu 等ICLR 2026 · 被引用 271 次
相关 Paper
- Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable RewardsShangyu Xing, Siyuan Wang, Chenyuan Yang, Xin-Yu Dai 等ICLR 2026 · 被引用 14 次
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-ForgetLiang Chen, Xueting Han, Qizhou Wang, Bo Han 等ICLR 2026 · 被引用 16 次
- ReLaX: Reasoning with Latent Exploration for Large Reasoning ModelsShimin Zhang, Xianwei Chen, Yufan Shen, Ziyuan Ye 等CVPR 2026 · 被引用 4 次
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy OptimizationHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu 等ICML 2026
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
