Scaling Laws for the Value of Individual Data Points in Machine Learning
Ian Connick Covert, Wenlong Ji, Tatsunori Hashimoto, James Zou
摘要
Recent works have shown that machine learning models improve at a predictable rate with the total amount of training data, leading to scaling laws that describe the relationship between error and dataset size. These scaling laws can help design a model's training dataset, but they typically take an aggregate view of the data by only considering the dataset's size. We introduce a new perspective by investigating scaling behavior for the value of individual data points: we find that a data point's contribution to model's performance shrinks predictably with the size of the dataset in a log-linear manner. Interestingly, there is significant variability in the scaling exponent among different data points, indicating that certain points are more valuable in small datasets while others are relatively more useful as a part of large datasets. We provide learning theory to support our scaling law, and we observe empirically that it holds across diverse model classes. We further propose a maximum likelihood estimator and an amortized estimator to efficiently learn the individualized scaling behaviors from a small number of noisy observations per data point. Using our estimators, we provide insights into factors that influence the scaling behavior of different data points. Finally, we demonstrate applications of the individualized scaling laws to data valuation and data subset selection. Overall, our work represents a first step towards understanding and utilizing scaling properties for the value of individual data points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Data-Efficient Selection via Grammatical Complexity in Continual Pre-training of Domain-Specific LLMsYizhou Ying, Geng Zhang, Cui Danxin, Chengyu Du 等EMNLP 2025
- (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep LearningMargaret Li, Sneha Kudugunta, Luke ZettlemoyerICLR 2025
- Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceJiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan 等ICLR 2025
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
相关 Paper
- Measuring the Effect of Training Data on Deep Learning Predictions via Randomized ExperimentsJinkun Lin, Anqi Zhang, Mathias Lécuyer, Jinyang Li 等ICML 2022 · 被引用 70 次
- Model Performance Scaling with Multiple Data SourcesTatsunori HashimotoICML 2021 · 被引用 38 次
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng 等ICML 2024 · 被引用 6 次
- How Much More Data Do I Need? Estimating Requirements for Downstream TasksRafid Mahmood, James Lucas, David Acuna, Daiqing Li 等CVPR 2022 · 被引用 21 次
- Improved Bayes Risk Can Yield Reduced Social Welfare Under CompetitionMeena Jagadeesan, Michael I. Jordan, Jacob Steinhardt, Nika HaghtalabNeurIPS 2023 · 被引用 20 次
