A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
Yuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth, James Y. Zou, Nan Jiang, Jiaqi Ma, Han Zhao
Abstract
Online reinforcement learning (RL) excels in complex, safety-critical domains but suffers from sample inefficiency, training instability, and limited interpretability. Data attribution provides a principled way to trace model behavior back to training samples, yet existing methods assume fixed datasets, which is violated in online RL where each experience both updates the policy and shapes future data collection. In this paper, we initiate the study of data attribution for online RL, focusing on the widely used Proximal Policy Optimization (PPO) algorithm. We start by establishing a local attribution framework, interpreting model checkpoints with respect to the records in the recent training buffer. We design two target functions, capturing agent action and cumulative return respectively, and measure each record's contribution through gradient similarity between its training loss and these targets. We demonstrate the power of this framework through three concrete applications: diagnosis of learning, temporal analysis of behavior formation, and targeted intervention during training. Leveraging this framework, we further propose an algorithm, iterative influence-based filtering (IIF), for online RL training that iteratively performs experience filtering to refine policy updates. Across standard RL benchmarks (classic control, navigation, locomotion) to RLHF for large language models, IIF reduces sample complexity, speeds up training, and achieves higher returns. Together, these results open a new direction for making online RL more interpretable, efficient, and effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76b4dd44-342e-4edc-a1a8-4c1c937f22d0Cited by top-tier papers6
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu et al.ICML 2026 · 12 citations
- Data Efficient RLVR via Off-Policy Influence GuidanceErle Zhu, Dazhi Jiang, Yuan Wang, Xujun Li et al.ACL 2026 · 7 citations
- Influence-based Online Experience Selection for Effective RLHFYifan Gong, Jing Yao, Xiting Wang, Xunlong Wang et al.ACL 2026
- GeoAlign: Geometric Rollout Curation for Robust LLM Reinforcement LearningTing Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen et al.ICML 2026
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
Builds on31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
Related papers
- OPPO: Accelerating PPO-based RLHF via Pipeline OverlapKaizhuo Yan, Yingjie Yu, Yifan Yu, Haizhong Zheng et al.ICLR 2026 · 4 citations
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee et al.ACL 2024 · 20 citations
- PS-PPO : Prefix-Sampling PPO for Critic-Free RLHFDoo Hwan Hwang, Kee-Eung KimICML 2026
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang et al.ICML 2026 · 22 citations
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
