Towards a statistical theory of data selection under weak supervision
Germain Kolossov, Andrea Montanari, Pulkit Tandon
摘要
Given a sample of size , it is often useful to select a subsample of smaller size to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the computational complexity of learning. We assume to be given unlabeled samples , and to be given access to a `surrogate model' that can predict labels better than random guessing. Our goal is to select a subset of the samples, to be denoted by , of size . We then acquire labels for this set and we use them to train a model via regularized empirical risk minimization. By using a mixture of numerical experiments on real and synthetic data, and mathematical derivations under low- and high- dimensional asymptotics, we show that: Data selection can be very effective, in particular beating training on the full sample in some cases; Certain popular choices in data selection methods (e.g. unbiased reweighted subsampling, or influence function-based subsampling) can be substantially suboptimal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Most Influential Subset Selection: Challenges, Promises, and BeyondYuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. MaNeurIPS 2024 · 被引用 39 次
- SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory MatchingYongmin Lee, Hye Won ChungICML 2024 · 被引用 26 次
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang 等NeurIPS 2025 · 被引用 22 次
- Sketchy Moment Matching: Toward Fast and Provable Data Selection for FinetuningYijun Dong, Viet Hoang Phan, Xiang Pan, Qi LeiNeurIPS 2024 · 被引用 9 次
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton 等ICLR 2025 · 被引用 6 次
它引用的顶会 Paper3
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He 等AAAI 2020 · 被引用 61 次
相关 Paper
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 被引用 30 次
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 被引用 25 次
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 被引用 54 次
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao 等VLDB 2022 · 被引用 38 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
