Understanding the detrimental class-level effects of data augmentation
Polina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt, Shanmukha Ramakrishna Vedantam, Hamed Firooz, Andrew Gordon Wilson
Abstract
Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average accuracy, recent studies have shown that its impact can be highly class dependent: achieving optimal average accuracy comes at the cost of significantly hurting individual class accuracy by as much as 20% on ImageNet. There has been little progress in resolving class-level accuracy drops due to a limited understanding of these effects. In this work, we present a framework for understanding how DA interacts with class-level learning dynamics. Using higherquality multi-label annotations on ImageNet, we systematically categorize the affected classes and find that the majority are inherently ambiguous, co-occur, or involve fine-grained distinctions, while DA controls the model's bias towards one of the closely related classes. While many of the previously reported performance drops are explained by multi-label annotations, our analysis of class confusions reveals other sources of accuracy degradation. We show that simple class-conditional augmentation strategies informed by our framework improve performance on the negatively affected classes. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b16a3485-46b8-46c6-b2bf-05c2987a50e6Cited by top-tier papers7
- TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based ModelsAndrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 19 citations
- What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable InsightsXin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang et al.NeurIPS 2024 · 19 citations
- LookHere: Vision Transformers with Directed Attention Generalize and ExtrapolateAnthony Fuller, Daniel G. Kyrollos, Yousef Yassin, James R. GreenNeurIPS 2024 · 10 citations
- Non-Asymptotic Analysis Of Data Augmentation For Precision Matrix EstimationLucas Morisset, Adrien Hardy, Alain DurmusNeurIPS 2025 · 2 citations
- No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth DataDaniel Cai, Randall BalestrieroICLR 2025
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- TrivialAugment: Tuning-free Yet State-of-the-Art Data AugmentationSamuel G. Müller, Frank HutterICCV 2021 · 384 citations
Related papers
- The Effects of Regularization and Data Augmentation are Class DependentRandall Balestriero, Léon Bottou, Yann LeCunNeurIPS 2022 · 124 citations
- Classes Are Not Equal: An Empirical Study on Image Recognition FairnessJiequan Cui, Beier Zhu, Xin Wen, Xiaojuan Qi et al.CVPR 2024
- Reducing Class-Wise Performance Disparity via Margin RegularizationBeier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun et al.ICLR 2026 · 1 citation
- Combining Ensembles and Data Augmentation Can Harm Your CalibrationYeming Wen, Ghassen Jerfel, Rafael Muller, Michael W. Dusenberry et al.ICLR 2021 · 72 citations
- Sharpen Focus: Learning With Attention Separability and ConsistencyLezi Wang, Ziyan Wu, Srikrishna Karanam, Kuan-Chuan Peng et al.ICCV 2019 · 37 citations
