Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Haizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura, Fan Lai, Jiawei Zhao, Beidi Chen
摘要
Reinforcement learning, such as PPO and GRPO, has powered recent breakthroughs in LLM reasoning. Scaling rollout to sample more prompts enables models to selectively use higher-quality data for training, which can stabilize RL training and improve model performance, but at the cost of significant computational overhead. In this paper, we first show that a substantial portion of this overhead can be avoided by skipping uninformative prompts before rollout. Our analysis of reward dynamics reveals a strong temporal consistency in prompt value: prompts that are uninformative in one epoch of training are likely to remain uninformative in near future epochs. Based on these insights, we propose GRESO (GRPO with Efficient Selective Rollout), an online, lightweight pre-rollout filtering algorithm that predicts and skips uninformative prompts using reward training dynamics. By evaluating GRESO on a broad range of math reasoning benchmarks and models, like Qwen2.5-Math-1.5B, DeepSeek-R1-Distill-Qwen-1.5B, Qwen2.5-Math-7B, Qwen2.5-14B, and Qwen2.5-32B, we show that GRESO achieves up to 2.4× wall-clock time speedup in rollout and up to 2.0× speedup in total training time without accuracy degradation. We make our code publicly available at GitHub 1 . 4.3M fewer rollouts 6.7 M fewer rollouts
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- The Art of Scaling Reinforcement Learning Compute for LLMsDevvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal 等ICLR 2026 · 被引用 95 次
- Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?Haizhong Zheng, Jiawei Zhao, Beidi ChenICLR 2026 · 被引用 57 次
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai 等ICLR 2026 · 被引用 55 次
- Prompt Curriculum Learning for Efficient LLM Post-TrainingZhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims 等ICLR 2026 · 被引用 44 次
- Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?Yun Qu, Qi Wang, Yixiu Mao, Vincent Tao Hu 等KDD 2026 · 被引用 33 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVRHaobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan 等ACL 2026 · 被引用 6 次
- Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning ModelsYixiu Mao, Yun Qu, Qi Wang, Heming Zou 等ICLR 2026 · 被引用 13 次
- Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning ModelsYun Qu, Qi Wang, Yixiu Mao, Heming Zou 等ICML 2026 · 被引用 7 次
- Unbiased Dynamic Pruning for Efficient Group-Based Policy OptimizationHaodong Zhu, Ren Yangyang, Yanjing Li, Mingbao Lin 等ICML 2026 · 被引用 2 次
