A Data-Augmentation Is Worth A Thousand Samples: Analytical Moments And Sampling-Free Training
Randall Balestriero, Ishan Misra, Yann LeCun
Abstract
Data-Augmentation (DA) is known to improve performance across tasks and datasets. We propose a method to theoretically analyze the effect of DA and study questions such as: how many augmented samples are needed to correctly estimate the information encoded by that DA? How does the augmentation policy impact the final parameters of a model? We derive several quantities in close-form, such as the expectation and variance of an image, loss, and model’s output under a given DA distribution. Up to our knowledge, we obtain the first explicit regularizer that corresponds to using DA during training for non-trivial transformations such as affine transformations, color jittering, or Gaussian blur. Those derivations open new avenues to quantify the benefits and limitations of DA. For example, given a loss at hand, we find that common DAs require tens of thousands of samples for the loss to be correctly estimated and for the model training to converge. We then show that for a training loss to have reduced variance under DA sampling, the model’s saliency map (gradient of the loss with respect to the model’s input) must align with the smallest eigenvector of the sample’s covariance matrix under the considered DA augmentation; this is exactly the quantity estimated and regularized by TangentProp. Those findings also hint at a possible explanation on why models tend to shift their focus from edges to textures when specific DAs are employed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 099afccc-004d-4d00-a7c9-351945db1ee8Cited by top-tier papers8
- No Representation Rules Them All in Category DiscoverySagar Vaze, Andrea Vedaldi, Andrew ZissermanNeurIPS 2023 · 79 citations
- Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self-Supervised LearningHugues Van Assel, Mark Ibrahim, Tommaso Biancalani, Aviv Regev et al.NeurIPS 2025 · 39 citations
- SF(DA)2: Source-free Domain Adaptation Through the Lens of Data AugmentationUiwon Hwang, Jonghyun Lee, Juhyeon Shin, Sungroh YoonICLR 2024 · 31 citations
- How Much Data Are Augmentations Worth? An Investigation into Scaling Laws, Invariance, and Implicit RegularizationJonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv et al.ICLR 2023 · 11 citations
- Revisiting Data Augmentation in Deep Reinforcement LearningJianshu Hu, Yunpeng Jiang, Paul WengICLR 2024 · 9 citations
Builds on5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
Related papers
- Tradeoffs in Data Augmentation: An Empirical StudyRaphael Gontijo Lopes, Sylvia J. Smullin, Ekin Dogus Cubuk, Ethan DyerICLR 2021 · 72 citations
- Data augmentation for deep learning based accelerated MRI reconstruction with limited dataZalan Fabian, Reinhard Heckel, Mahdi SoltanolkotabiICML 2021 · 60 citations
- First-Order Manifold Data Augmentation for Regression LearningIlya Kaufman, Omri AzencotICML 2024 · 6 citations
- KeepAugment: A Simple Information-Preserving Data Augmentation ApproachChengyue Gong, Dilin Wang, Meng Li, Vikas Chandra et al.CVPR 2021
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
