Geometric Median (GM) Matching for Robust k-Subset Selection from Noisy Data
Anish Acharya, Sujay Sanghavi, Alex Dimakis, Inderjit S. Dhillon
Abstract
Data pruning -the combinatorial task of selecting a small and representative subset from a large dataset, is crucial for mitigating the enormous computational costs associated with training datahungry modern deep learning models at scale. Since large-scale data collections are invariably noisy, developing data pruning strategies that remain robust even in the presence of corruption is critical in practice. Existing data pruning methods often fail under high corruption rates due to their reliance on empirical mean estimation, which is highly sensitive to outliers. In response, this work proposes Geometric Median (GM) Matching, a novel k-subset selection strategy that leverages the GM, a robust estimator with an optimal breakdown point of 1/2; to enhance resilience against noisy data. Our method iteratively selects a ksubset such that the mean of the subset approximates the GM of the (potentially) noisy dataset, ensuring robustness even under arbitrary corruption. We provide theoretical guarantees, showing that GM MATCHING enjoys an improved O(1/k) convergence rate, outperforming O(1/ √ k) scaling of uniform sampling, even under arbitrary corruption. Extensive experiments across image classification and image generation tasks demonstrate that GM MATCHING consistently outperforms existing pruning approaches, particularly in high-corruption settings; making it a strong baseline for robust data pruning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e393c551-036f-4ffe-9670-387e4fd5d4a2Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset SelectionWeiwei Xiao, Yongyong Chen, Qiben Shan, Yaowei Wang et al.AAAI 2024 · 14 citations
- Robust Data Pruning under Label Noise via Maximizing Re-labeling AccuracyDongmin Park, Seola Choi, Doyoung Kim, Hwanjun Song et al.NeurIPS 2023 · 42 citations
- Procore: Robust Core-Set Selection Via Pareto Multi-Dimensional Optimization From Noisy DataXiaoou Ding, Hongbin Hu, Songnan Jiang, Muyun Zhou et al.ICDE 2026
- DRoP: Distributionally Robust Data PruningArtem M. Vysogorets, Kartik Ahuja, Julia KempeICLR 2025
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu et al.ICLR 2023 · 21 citations
