Towards a statistical theory of data selection under weak supervision
Germain Kolossov, Andrea Montanari, Pulkit Tandon
Abstract
Given a sample of size , it is often useful to select a subsample of smaller size to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the computational complexity of learning. We assume to be given unlabeled samples , and to be given access to a `surrogate model' that can predict labels better than random guessing. Our goal is to select a subset of the samples, to be denoted by , of size . We then acquire labels for this set and we use them to train a model via regularized empirical risk minimization. By using a mixture of numerical experiments on real and synthetic data, and mathematical derivations under low- and high- dimensional asymptotics, we show that: Data selection can be very effective, in particular beating training on the full sample in some cases; Certain popular choices in data selection methods (e.g. unbiased reweighted subsampling, or influence function-based subsampling) can be substantially suboptimal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 796db503-5c2a-44d0-bfb0-59b123d09e75Cited by top-tier papers21
- Most Influential Subset Selection: Challenges, Promises, and BeyondYuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. MaNeurIPS 2024 · 39 citations
- SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory MatchingYongmin Lee, Hye Won ChungICML 2024 · 26 citations
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang et al.NeurIPS 2025 · 22 citations
- Sketchy Moment Matching: Toward Fast and Provable Data Selection for FinetuningYijun Dong, Viet Hoang Phan, Xiang Pan, Qi LeiNeurIPS 2024 · 9 citations
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton et al.ICLR 2025 · 6 citations
Builds on3
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He et al.AAAI 2020 · 61 citations
Related papers
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 25 citations
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 54 citations
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao et al.VLDB 2022 · 38 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
