Provable Benefit of Cutout and CutMix for Feature Learning
Junsoo Oh, Chulhee Yun
Abstract
Patch-level data augmentation techniques such as Cutout and CutMix have demonstrated significant efficacy in enhancing the performance of vision tasks. However, a comprehensive theoretical understanding of these methods remains elusive. In this paper, we study two-layer neural networks trained using three distinct methods: vanilla training without augmentation, Cutout training, and CutMix training. Our analysis focuses on a feature-noise data model, which consists of several label-dependent features of varying rarity and label-independent noises of differing strengths. Our theorems demonstrate that Cutout training can learn low-frequency features that vanilla training cannot, while CutMix training can learn even rarer features that Cutout cannot capture. From this, we establish that CutMix yields the highest test accuracy among the three. Our novel analysis reveals that CutMix training makes the network learn all features and noise vectors"evenly"regardless of the rarity and strength, which provides an interesting insight into understanding patch-level augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c92f0e3d-fee1-4320-a52a-cd52d2a0f188Cited by top-tier papers3
- DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of PlasticityBaekrok Shin, Junsoo Oh, Hanseul Cho, Chulhee YunNeurIPS 2024 · 8 citations
- From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature LearningJunsoo Oh, Jerry Song, Chulhee YunNeurIPS 2025 · 5 citations
- Toward Understanding Adversarial Distillation: Why Robust Teachers FailHongsin Lee, Hye Won ChungICML 2026
Builds on28
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 457 citations
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani et al.ICLR 2021 · 294 citations
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
Related papers
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 36 citations
- A Unified Analysis of Mixed Sample Data Augmentation: A Loss Function PerspectiveChanwoo Park, Sangdoo Yun, Sanghyuk ChunNeurIPS 2022 · 43 citations
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 13 citations
- StyleMix: Separating Content and Style for Enhanced Data AugmentationMinui Hong, Jinwoo Choi, Gunhee KimCVPR 2021
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
