DRoP: Distributionally Robust Data Pruning
Artem M. Vysogorets, Kartik Ahuja, Julia Kempe
摘要
In the era of exceptionally data-hungry models, careful selection of the training data is essential to mitigate the extensive costs of deep learning. Data pruning offers a solution by removing redundant or uninformative samples from the dataset, which yields faster convergence and improved neural scaling laws. However, little is known about its impact on classification bias of the trained models. We conduct the first systematic study of this effect and reveal that existing data pruning algorithms can produce highly biased classifiers. We present theoretical analysis of the classification risk in a mixture of Gaussians to argue that choosing appropriate class pruning ratios, coupled with random pruning within classes has potential to improve worst-class performance. We thus propose DRoP, a distributionally robust approach to pruning and empirically demonstrate its performance on standard computer vision benchmarks. In sharp contrast to existing algorithms, our proposed method continues improving distributional robustness at a tolerable drop of average performance as we prune more from the datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
相关 Paper
- Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction UncertaintyYeseul Cho, Baekrok Shin, Changmin Kang, Chulhee YunICML 2025
- Robust Data Pruning under Label Noise via Maximizing Re-labeling AccuracyDongmin Park, Seola Choi, Doyoung Kim, Hwanjun Song 等NeurIPS 2023 · 被引用 42 次
- Bias in Pruned Vision Models: In-Depth Analysis and CountermeasuresEugenia Iofinova, Alexandra Peste, Dan AlistarhCVPR 2023
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin 等NeurIPS 2022 · 被引用 36 次
- Recall Distortion in Neural Network Pruning and the Undecayed Pruning AlgorithmAidan Good, Jiaqi Lin, Xin Yu, Hannah Sieg 等NeurIPS 2022 · 被引用 15 次
