Data Efficient RLVR via Off-Policy Influence Guidance
Erle Zhu, Dazhi Jiang, Yuan Wang, Xujun Li, Jiale Cheng, Yuxian Gu, Yilin Niu, Aohan Zeng, Jie Tang, Minlie Huang, Hongning Wang
摘要
Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large language models (LLMs). Current data selection methods are largely heuristicbased, lacking theoretical guarantees and generalizability. This work proposes a theoreticallygrounded approach using influence functions to estimate the contribution of each data point to the learning objective. To overcome the prohibitive computational cost of policy rollouts required for online influence estimation, we introduce an off-policy influence estimation method that efficiently approximates data influence using pre-collected offline trajectories. Furthermore, to manage the high-dimensional gradients of LLMs, we employ sparse random projection to reduce dimensionality and improve storage and computation efficiency. Leveraging these techniques, we develop Curriculum RL with Off-Policy Influence guidance (CROPI), a multi-stage RL framework that iteratively selects the most influential data for the current policy. Experiments on models up to 7B parameters demonstrate that CROPI significantly accelerates training. On a 1.5B model, it achieves a 2.66× step-level acceleration while using only 10% of the data per stage compared to full-dataset training. Our results highlight the substantial potential of influence-based data selection for efficient RLVR. Code is available at https://github.com/thu-coai/CROPI .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- DsDm: Model-Aware Dataset Selection with DatamodelsLogan Engstrom, Axel Feldmann, Aleksander MadryICML 2024 · 被引用 105 次
相关 Paper
- Towards High Data Efficiency in Reinforcement Learning with Verifiable RewardXinyu Tang, Zhenduo Zhang, Yurou Liu, Xin Zhao 等ICLR 2026 · 被引用 18 次
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang 等ACL 2026
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model ReasoningYukun Chen, Jiaming Li, Longze Chen, Ze Gong 等ICML 2026 · 被引用 5 次
- CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMsYongcheng Zeng, Zexu Sun, Bokai Ji, Erxue Min 等ICLR 2026 · 被引用 18 次
- Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement LearningYuheng Zha, Kun Zhou, Yujia Wu, Yushu Wang 等AAAI 2026
