Data Shapley in One Training Run
Jiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi Jia
摘要
Data Shapley offers a principled framework for attributing the contribution of data within machine learning contexts. However, the traditional notion of Data Shapley requires re-training models on various data subsets, which becomes computationally infeasible for large-scale models. Additionally, this retraining-based definition cannot evaluate the contribution of data for a specific model training run, which may often be of interest in practice. This paper introduces a novel concept, In-Run Data Shapley, which eliminates the need for model retraining and is specifically designed for assessing data contribution for a particular model of interest. In-Run Data Shapley calculates the Shapley value for each gradient update iteration and accumulates these values throughout the training process. We present several techniques that allow the efficient scaling of In-Run Data Shapley to the size of foundation models. In its most optimized implementation, our method adds negligible runtime overhead compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal 等NeurIPS 2024 · 被引用 91 次
- Regression-adjusted Monte Carlo Estimators for Shapley Values and Probabilistic ValuesR. Teal Witter, Yurong Liu, Christopher MuscoNeurIPS 2025 · 被引用 22 次
- DataRater: Meta-Learned Dataset CurationDan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf 等NeurIPS 2025 · 被引用 17 次
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu 等ICML 2026 · 被引用 12 次
它引用的顶会 Paper21
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
相关 Paper
- An Efficient Framework for Crediting Data Contributors of Diffusion ModelsMingyu Lu, Chris Lin, Chanwoo Kim, Su-In LeeICLR 2025
- SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) ModelsMingYu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim 等ICML 2026
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan 等KDD 2024 · 被引用 3 次
- Ripple Shapley: Data Influence Attribution in One Federated Training RunDewen Zeng, Wenlong Tian, Haozhao Wang, Jianfeng Lu 等AAAI 2026 · 被引用 1 次
- CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationStephanie Schoch, Haifeng Xu, Yangfeng JiNeurIPS 2022 · 被引用 56 次
