Lune

ACL2026顶会

I²B-LPO: Latent Policy Optimization via Iterative Information Bottleneck

Huilin Deng, Hongchen Luo, Yue Zhu, Long Li, Zhuoyue Chen, Xinghao Zhao, Ming Li, Chuyang Zhao, Jihai Zhang, Mengchang Wang, Yang Cao, Yu Kang

2026年份

摘要

Despite recent advances in Reinforcement learning with verifiable rewards (RLVR) for large language model (LLM) reasoning, most methods suffer from exploration collapse, as the semantic homogeneity of random rollouts traps models in narrow, over-optimized behaviors. Existing methods leverage policy entropy to encourage exploration, but face inherent limitations: global entropy regularization is susceptible to reward hacking, inducing meaningless verbosity, whereas local token-selective updates struggle with the strong inductive bias of pre-trained models. To this end, we propose Latent Policy Optimization via Iterative Information Bottleneck (I 2 B-LPO), which shifts from statistical perturbation of token distributions to topological branching of reasoning trajectories. I 2 B-LPO triggers latent branching at high-entropy states to diversify reasoning trajectories and applies the Information Bottleneck as a trajectory filter and self-reward to ensure concise and informative exploration. Empirical results on four mathematical benchmarks demonstrate that I 2 B-LPO achieves stateof-the-art performance, with margins of up to 5.3% in accuracy and 7.4% in diversity metrics. Code is available at https://github. com/denghuilin-cyber/IIB-LPO . Flatten across the whole vocabulary H z1 H … Critical position Quantile borderline Tokens Tokens (a) Entropy Regularization Method (b) Token-selective Method R: Basically, to actually find the minimum value, well, in the context of optimzation, first, we should differentiate … R: To find the minimum value of the function, let us differentiate the function f(x) with respect to x. (c) I² B-LPO(Ours) … Let us diff... the … Let us diff... the … first we should … … first we should … z2 zi zi zi zi zi zi zi ... Tt Tt+1 ... Tt+1 Tt Tt+n R 1 :we can differentiate the function f(x) with respect to x: R 1 :we can differentiate the function f(x) with respect to x: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R 2 : we analyze the structure. Note that f(x) can be rewritten as a sum of distances: R1 R2 To find the minimum value, first, Sharpen on critical position V'

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖