Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and Beyond
Kyriakos Axiotis, Vincent Cohen-Addad, Monika Henzinger, Sammy Jerome, Vahab Mirrokni, David Saulpic, David P. Woodruff, Michael Wunder
Abstract
We study the data selection problem, whose aim is to select a small representative subset of data that can be used to efficiently train a machine learning model. We present a new data selection approach based on -means clustering and sensitivity sampling. Assuming access to an embedding representation of the data with respect to which the model loss is Hölder continuous, our approach provably allows selecting a set of ``typical'' elements whose average loss corresponds to the average loss of the whole dataset, up to a multiplicative factor and an additive , where represents the -means cost for the input embeddings and is the Hölder constant. We furthermore demonstrate the performance and scalability of our approach on fine-tuning foundation models and show that it outperforms state-of-the-art methods. We also show how it can be applied on linear regression, leading to a new sampling strategy that surprisingly matches the performances of leverage score sampling, while being conceptually simpler and more scalable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38573c43-e54a-4b82-b828-3410bf09f6e7Cited by top-tier papers7
- Efficient Data Selection at Scale via Influence DistillationMahdi Nikdan, Vincent Cohen-Addad, Dan Alistarh, Vahab MirrokniNeurIPS 2025 · 15 citations
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang et al.ICML 2026 · 13 citations
- Sketchy Moment Matching: Toward Fast and Provable Data Selection for FinetuningYijun Dong, Viet Hoang Phan, Xiang Pan, Qi LeiNeurIPS 2024 · 9 citations
- Train on Validation (ToV): Fast data selection with applications to fine-tuningAyush Jain, Andrea Montanari, Eren SasogluICLR 2026 · 4 citations
- FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information GainRohan Deb, Kiran Koshy Thekumparampil, Kousha Kalantari, Gaurush Hiranandani et al.ICML 2025
Builds on12
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas et al.NeurIPS 2021 · 220 citations
- Data-Independent Neural Pruning via CoresetsBen Mussay, Margarita Osadchy, Vladimir Braverman, Samson Zhou et al.ICLR 2020 · 65 citations
- Coresets for Near-Convex FunctionsMurad Tukan, Alaa Maalouf, Dan FeldmanNeurIPS 2020 · 49 citations
- Improved Coresets for Euclidean k-MeansVincent Cohen-Addad, Kasper Green Larsen, David Saulpic, Chris Schwiegelshohn et al.NeurIPS 2022 · 47 citations
- Fast and Accurate -means++ via Rejection SamplingVincent Cohen-Addad, Silvio Lattanzi, Ashkan Norouzi-Fard, Christian Sohler et al.NeurIPS 2020 · 32 citations
Related papers
- Active Learning with Low-Rank Structure for Data SelectionVincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan et al.ICML 2026
- Sensitivity Sampling for k-Means: Worst Case and Stability Optimal Coreset BoundsNikhil Bansal, Vincent Cohen-Addad, Milind Prabhu, David Saulpic et al.FOCS 2024 · 2 citations
- Optimal bounds for ℓp sensitivity sampling via ℓ2 augmentationAlexander Munteanu, Simon OmlorICML 2024 · 6 citations
- Sharper Bounds for ℓp Sensitivity SamplingDavid P. Woodruff, Taisuke YasudaICML 2023 · 8 citations
- Near-optimal Coresets for Robust ClusteringLingxiao Huang, Shaofeng H.-C. Jiang, Jianing Lou, Xuan WuICLR 2023 · 1 citation
