Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
Yifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen, Zhendong Wang, Mingyuan Zhou, Huan Zhang
Abstract
Reinforcement learning (RL) has become an effective approach for fine-tuning large language models (LLMs), particularly to enhance their reasoning capabilities. However, RL fine-tuning remains highly resource-intensive, and existing work has largely overlooked the problem of data efficiency. In this paper, we propose two techniques to improve data efficiency in LLM RL fine-tuning: difficultytargeted online data selection and rollout replay. We introduce the notion of adaptive difficulty to guide online data selection, prioritizing questions of moderate difficulty that are more likely to yield informative learning signals. To estimate adaptive difficulty efficiently, we develop an attention-based framework that requires rollouts for only a small reference set of questions. The adaptive difficulty of the remaining questions is then estimated based on their similarity to this set. To further reduce rollout cost, we introduce a rollout replay mechanism inspired by experience replay in traditional RL. This technique reuses recent rollouts, lowering per-step computation while maintaining stable updates. Experiments across 6 LLM-dataset combinations show that our method reduces RL fine-tuning time by 23% to 62% while reaching the same level of performance as the original GRPO algorithm. Our code repository is available at https://github.com/ASTRAL-Group/data-efficient-llm-rl/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8d2187e-e53f-4f71-97f4-cd6ec6ed0fa7Cited by top-tier papers19
- Prompt Curriculum Learning for Efficient LLM Post-TrainingZhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims et al.ICLR 2026 · 44 citations
- Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable RewardsHieu Trung Nguyen, Bao Nguyen, Wenao Ma, Yuzhi Zhao et al.ICLR 2026 · 19 citations
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement FinetuningQianli Shen, Daoyuan Chen, Yilun Huang, Zhenqing Ling et al.ICLR 2026 · 15 citations
- Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuningSirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun et al.ICLR 2026 · 9 citations
- Data Efficient RLVR via Off-Policy Influence GuidanceErle Zhu, Dazhi Jiang, Yuan Wang, Xujun Li et al.ACL 2026 · 7 citations
Builds on14
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang et al.NeurIPS 2025 · 533 citations
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
Related papers
- PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy OptimizationXinxin Zhu, Ying He, Haowen Hou, Ruichong Zhang et al.AAAI 2026
- Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsMengqi Liao, Xiangyu Xi, Ruinian Chen, Jia Leng et al.EMNLP 2025 · 16 citations
- Knapsack RL: Compute-Efficient Reinforcement Learning via Heterogeneous Rollout AllocationZiniu Li, Congliang Chen, Tianyun Yang, Tian Ding et al.ICML 2026
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic ShuffleLinghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju et al.ICLR 2026 · 18 citations
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen et al.ICML 2026
