Geometric Median (GM) Matching for Robust k-Subset Selection from Noisy Data
Anish Acharya, Sujay Sanghavi, Alex Dimakis, Inderjit S. Dhillon
摘要
Data pruning -the combinatorial task of selecting a small and representative subset from a large dataset, is crucial for mitigating the enormous computational costs associated with training datahungry modern deep learning models at scale. Since large-scale data collections are invariably noisy, developing data pruning strategies that remain robust even in the presence of corruption is critical in practice. Existing data pruning methods often fail under high corruption rates due to their reliance on empirical mean estimation, which is highly sensitive to outliers. In response, this work proposes Geometric Median (GM) Matching, a novel k-subset selection strategy that leverages the GM, a robust estimator with an optimal breakdown point of 1/2; to enhance resilience against noisy data. Our method iteratively selects a ksubset such that the mean of the subset approximates the GM of the (potentially) noisy dataset, ensuring robustness even under arbitrary corruption. We provide theoretical guarantees, showing that GM MATCHING enjoys an improved O(1/k) convergence rate, outperforming O(1/ √ k) scaling of uniform sampling, even under arbitrary corruption. Extensive experiments across image classification and image generation tasks demonstrate that GM MATCHING consistently outperforms existing pruning approaches, particularly in high-corruption settings; making it a strong baseline for robust data pruning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset SelectionWeiwei Xiao, Yongyong Chen, Qiben Shan, Yaowei Wang 等AAAI 2024 · 被引用 14 次
- Robust Data Pruning under Label Noise via Maximizing Re-labeling AccuracyDongmin Park, Seola Choi, Doyoung Kim, Hwanjun Song 等NeurIPS 2023 · 被引用 42 次
- Procore: Robust Core-Set Selection Via Pareto Multi-Dimensional Optimization From Noisy DataXiaoou Ding, Hongbin Hu, Songnan Jiang, Muyun Zhou 等ICDE 2026
- DRoP: Distributionally Robust Data PruningArtem M. Vysogorets, Kartik Ahuja, Julia KempeICLR 2025
- Dataset Pruning: Reducing Training Data by Examining Generalization InfluenceShuo Yang, Zeke Xie, Hanyu Peng, Min Xu 等ICLR 2023 · 被引用 21 次
