Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
Haizhong Zheng, Jiawei Zhao, Beidi Chen
摘要
Reinforcement learning has been central to recent advances in large language model reasoning, but most algorithms rely on on-policy training that demands fresh rollouts at every update, limiting efficiency and scalability. Asynchronous RL systems alleviate this by decoupling rollout generation from training, yet their effectiveness hinges on tolerating large staleness in rollout data, a setting where existing methods either degrade in performance or collapse. We revisit this challenge and uncover a prosperity-before-collapse phenomenon: stale data can be as informative as on-policy data if exploited properly. Building on this insight, we introduce M2PO (Second-Moment Trust Policy Optimization), which constrains the second moment of importance weights to suppress only extreme outliers while preserving informative updates. Notably, M2PO sharply reduces the fraction of clipped tokens under high staleness (from 1.22% to 0.06% over training), precisely masking high-variance tokens while maintaining stable optimization. Extensive evaluation across six model scales (from 1.7B to 32B) and eight math reasoning benchmarks and one coding benchmarks shows that M2PO delivers stable off-policy training even with data stale by at least 256 model updates and matches on-policy performance. Our code is available at https://github.com/Infini-AI-Lab/M2PO/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingBrian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain 等NeurIPS 2025 · 被引用 34 次
- RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMsYongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu 等NSDI 2026 · 被引用 4 次
- Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMsLuke Huang, Zhuoyang Zhang, Qinghao Hu, Shang Yang 等ICML 2026 · 被引用 3 次
- : A Generalist Value Model for Any Policy at State ZeroYi-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun 等ICML 2026 · 被引用 3 次
- QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuningDoyeon Lee, Eunyi Lyou, Hyunsoo Cho, Soo Kyung Kim 等ICML 2026
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
相关 Paper
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement LearningZhenpeng Su, Leiyu Pan, Minxuan Lv, Yuntao Li 等ACL 2026 · 被引用 21 次
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM ReasoningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalICLR 2026 · 被引用 7 次
- BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive ClippingZhiheng Xi, Xin Guo, Yang Nan, Enyu Zhou 等ICLR 2026 · 被引用 54 次
