Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Yingru Li, Jiacai Liu, Jiawei Xu, Yuxuan Tong, Ziniu Li, Baoxiang Wang
Abstract
Policy gradient methods for Large Language Models optimize a policy π θ via a surrogate objective computed from samples of a rollout policy π roll . However, modern LLM-RL pipelines suffer from unavoidable implementation divergences-backend discrepancies, Mixtureof-Experts routing discontinuities, and distributed training staleness-causing off-policy mismatch (π roll ̸ = π θ ) and approximation errors between the surrogate and the true objective. We demonstrate that classical trust region bounds on this error scale as O(T 2 ) with sequence length T , rendering them vacuous for long-horizon tasks. To address this, we derive a family of bounds-both KL-based and TV-based-including a Pinsker-Marginal bound (O(T 3/2 )), a Mixed bound (O(T )), and an Adaptive bound that strictly generalizes the Pinsker-Marginal bound via per-position importance-ratio decomposition. Taking the minimum over all bounds yields the tightest known guarantee across all divergence regimes. Crucially, all bounds depend on the maximum token-level divergence D tok,max KL (or D tok,max TV ), a sequence-level quantity that cannot be controlled by token-independent methods like PPO clipping. We propose Trust Region Masking (TRM), which masks entire sequences violating the trust region, enabling the first non-vacuous monotonic improvement guarantees for long-horizon LLM-RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference MismatchZiyang Zhang, Xinheng Ding, Jiayi Yuan, Rixin Liu et al.ICML 2026 · 13 citations
- The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RLYingru Li, Jiawei Xu, Ziniu Li, Jiacai Liu et al.ICML 2026 · 4 citations
Builds on3
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- TROLL: Trust Regions Improve Reinforcement Learning for Large Language ModelsPhilipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto et al.ICLR 2026 · 8 citations
- BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement LearningYuan Li, Bo Wang, Yufei Gao, Yuqian Yao et al.ICML 2026 · 2 citations
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang et al.ICML 2026 · 22 citations
- MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM ReasoningXiaoliang Fu, Jiaye Lin, Yangyi Fang, Binbin Zheng et al.ACL 2026 · 12 citations
- From łog π to π: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient WeightXiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu et al.ACL 2026 · 5 citations
