Lune

ICML2026顶会

Trust Region Masking for Long-Horizon LLM Reinforcement Learning

Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Baoxiang Wang

2026年份
2顶会引用

摘要

Policy gradient methods for Large Language Models optimize a policy π θ via a surrogate objective computed from samples of a rollout policy π roll . However, modern LLM-RL pipelines suffer from unavoidable implementation divergences-backend discrepancies, Mixtureof-Experts routing discontinuities, and distributed training staleness-causing off-policy mismatch (π roll ̸ = π θ ) and approximation errors between the surrogate and the true objective. We demonstrate that classical trust region bounds on this error scale as O(T 2 ) with sequence length T , rendering them vacuous for long-horizon tasks. To address this, we derive a family of bounds-both KL-based and TV-based-including a Pinsker-Marginal bound (O(T 3/2 )), a Mixed bound (O(T )), and an Adaptive bound that strictly generalizes the Pinsker-Marginal bound via per-position importance-ratio decomposition. Taking the minimum over all bounds yields the tightest known guarantee across all divergence regimes. Crucially, all bounds depend on the maximum token-level divergence D tok,max KL (or D tok,max TV ), a sequence-level quantity that cannot be controlled by token-independent methods like PPO clipping. We propose Trust Region Masking (TRM), which masks entire sequences violating the trust region, enabling the first non-vacuous monotonic improvement guarantees for long-horizon LLM-RL.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper3

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖