ACL2026

I²B-LPO: Latent Policy Optimization via Iterative Information Bottleneck

Huilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming Li, Chuyang Zhao, Jihai Zhang, Mengchang Wang, Yang Cao, Yu Kang

Abstract

Despite recent advances in Reinforcement learning with verifiable rewards (RLVR) for large language model (LLM) reasoning, most methods suffer from exploration collapse, as the semantic homogeneity of random rollouts traps models in narrow, over-optimized behaviors. Existing methods leverage policy entropy to encourage exploration, but face inherent limitations: global entropy regularization is susceptible to reward hacking, inducing meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To this end, we propose Latent Policy Optimization via Iterative Information Bottleneck (I 2 B-LPO), which shifts from statistical perturbation of token distributions to topological branching of reasoning trajectories. I 2 B-LPO triggers latent branching at high-entropy states to diversify reasoning trajectories and applies the Information Bottleneck as a trajectory filter and self-reward to ensure concise and informative exploration. Empirical results on four mathematical benchmarks demonstrate that I 2 B-LPO achieves stateof-the-art performance, with margins of up to 5.3% in accuracy and 7.4% in diversity metrics. Code is available at https://github. com/denghuilin-cyber/IIB-LPO . Flatten across the whole vocabulary H z1 H … Critical position Quantile borderline Tokens Tokens (a) Entropy Regularization Method (b) Token-selective Method R: Basically, to actually find the minimum value, well, in the context of optimzation, first, we should differentiate … R: To find the minimum value of the function, let us differentiate the function f(x) with respect to x. (c) I² B-LPO(Ours) … Let us diff... the … Let us diff... the … first we should … … first we should … z2 zi zi zi zi zi zi zi ... Tt Tt+1 ... Tt+1 Tt Tt+n R 1 :we can differentiate the function f(x) with respect to x: R 1 :we can differentiate the function f(x) with respect to x: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R1 R2 To find the minimum value, first, Sharpen on critical position V'