On the Generalization Effects of Linear Transformations in Data Augmentation
Sen Wu, Hongyang R. Zhang, Gregory Valiant, Christopher Ré
Abstract
Data augmentation is a powerful technique to improve performance in applications such as image and text classification tasks. Yet, there is little rigorous understanding of why and how various augmentations work. In this work, we consider a family of linear transformations and study their effects on the ridge estimator in an over-parametrized linear regression setting. First, we show that transformations that preserve the labels of the data can improve estimation by enlarging the span of the training data. Second, we show that transformations that mix data can improve estimation by playing a regularization effect. Finally, we validate our theoretical insights on MNIST. Based on the insights, we propose an augmentation scheme that searches over the space of transformations by how uncertain the model is about the transformed data. We validate our proposed scheme on image and text datasets. For example, our method outperforms random sampling methods by 1.24% on CIFAR-100 using Wide-ResNet-28-10. Furthermore, we achieve comparable accuracy to the SoTA Adversarial AutoAugment on CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- On Interaction Between Augmentations and Corruptions in Natural Corruption RobustnessEric Mintun, Alexander Kirillov, Saining XieNeurIPS 2021 · 138 citations
- Toward Understanding Generative Data AugmentationChenyu Zheng, Guoqiang Wu, Chongxuan LiNeurIPS 2023 · 51 citations
- Direct Differentiable Augmentation SearchAoming Liu, Zehao Huang, Zhiwu Huang, Naiyan WangICCV 2021 · 47 citations
- Noisy Feature MixupSoon Hoe Lim, N. Benjamin Erichson, Francisco Utrera, Winnie Xu et al.ICLR 2022 · 43 citations
- Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization GuaranteesHaotian Ju, Dongyue Li, Hongyang R. ZhangICML 2022 · 41 citations
Builds on2
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Understanding and Mitigating the Tradeoff between Robustness and AccuracyAditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi et al.ICML 2020 · 252 citations
Related papers
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani et al.ICLR 2021 · 294 citations
- Regularization properties of adversarially-trained linear regressionAntônio H. Ribeiro, Dave Zachariah, Francis R. Bach, Thomas B. SchönNeurIPS 2023 · 23 citations
- AutoDO: Robust AutoAugment for Biased Data With Label Noise via Scalable Probabilistic Implicit DifferentiationDenis A. Gudovskiy, Luca Rigazio, Shun Ishizaka, Kazuki Kozuka et al.CVPR 2021
- GradAug: A New Regularization Method for Deep Neural NetworksTaojiannan Yang, Sijie Zhu, Chen ChenNeurIPS 2020 · 43 citations
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 124 citations
