Data Augmentation as Feature Manipulation
Ruoqi Shen, Sébastien Bubeck, Suriya Gunasekar
Abstract
Data augmentation is a cornerstone of the machine learning pipeline, yet its theoretical underpinnings remain unclear. Is it merely a way to artificially augment the data set size? Or is it about encouraging the model to satisfy certain invariance? In this work we consider another angle, and we study the effect of data augmentation on the dynamic of the learning process. We find that data augmentation can alter the relative importance of various features, effectively making certain informative but hard to learn features more likely to be captured in the learning process. Importantly, we show that this effect is more pronounced for non-linear models, such as neural networks. Our main contribution is a detailed analysis of data augmentation on the learning dynamic for a two layer convolutional neural network in the recently proposed multi-view data model by Allen-Zhu and Li [2020b]. We complement this analysis with further experimental evidence that data augmentation can be viewed as feature manipulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Understanding and Improving Feature Learning for Out-of-Distribution GeneralizationYongqiang Chen, Wei Huang, Kaiwen Zhou, Yatao Bian et al.NeurIPS 2023 · 49 citations
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 36 citations
- ModelDiff: A Framework for Comparing Learning AlgorithmsHarshay Shah, Sung Min Park, Andrew Ilyas, Aleksander MadryICML 2023 · 36 citations
- Understanding Convergence and Generalization in Federated Learning through Feature Learning TheoryWei Huang, Ye Shi, Zhongyi Cai, Taiji SuzukiICLR 2024 · 17 citations
- Optimizing Data Acquisition to Enhance Machine Learning PerformanceTingting Wang, Shixun Huang, Zhifeng Bao, J. Shane Culpepper et al.VLDB 2024 · 13 citations
Builds on5
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- On the Generalization Effects of Linear Transformations in Data AugmentationSen Wu, Hongyang R. Zhang, Gregory Valiant, Christopher RéICML 2020 · 92 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- How Data Augmentation affects Optimization for Linear RegressionBoris Hanin, Yi SunNeurIPS 2021 · 22 citations
Related papers
- Towards Understanding Catastrophic Forgetting in Two-layer Convolutional Neural NetworksBoqi Li, Youjun Wang, Weiwei LiuICML 2025
- MetAug: Contrastive Learning via Meta Feature AugmentationJiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su et al.ICML 2022 · 32 citations
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 13 citations
- How Much Data Are Augmentations Worth? An Investigation into Scaling Laws, Invariance, and Implicit RegularizationJonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv et al.ICLR 2023 · 11 citations
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin et al.NeurIPS 2022 · 56 citations
