Provably Learning Diverse Features in Multi-View Data with Midpoint Mixup
Muthu Chidambaram, Xiang Wang, Chenwei Wu, Rong Ge
Abstract
Mixup is a data augmentation technique that relies on training using random convex combinations of data points and their labels. In recent years, Mixup has become a standard primitive used in the training of state-of-the-art image classification models due to its demonstrated benefits over empirical risk minimization with regards to generalization and robustness. In this work, we try to explain some of this success from a feature learning perspective. We focus our attention on classification problems in which each class may have multiple associated features (or views) that can be used to predict the class correctly. Our main theoretical results demonstrate that, for a non-trivial class of data distributions with two features per class, training a 2-layer convolutional network using empirical risk minimization can lead to learning only one feature for almost all classes while training with a specific instantiation of Mixup succeeds in learning both features for every class. We also show empirically that these theoretical insights extend to the practical settings of image benchmarks modified to have multiple features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23d1a7d1-0efe-4788-b7a4-0ef0fe70b748Cited by top-tier papers14
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 36 citations
- Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and MethodBikang Pan, Wei Huang, Ye ShiNeurIPS 2024 · 28 citations
- Understanding Convergence and Generalization in Federated Learning through Feature Learning TheoryWei Huang, Ye Shi, Zhongyi Cai, Taiji SuzukiICLR 2024 · 17 citations
- Provable Benefit of Cutout and CutMix for Feature LearningJunsoo Oh, Chulhee YunNeurIPS 2024 · 11 citations
- Pushing Boundaries: Mixup's Influence on Neural CollapseQuinn LeBlanc Fisher, Haoming Meng, Vardan PapyanICLR 2024 · 8 citations
Builds on13
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 457 citations
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
- Co-Mixup: Saliency Guided Joint Mixup with Supermodular DiversityJang-Hyun Kim, Wonho Choo, Hosan Jeong, Hyun Oh SongICLR 2021 · 207 citations
- InstaHide: Instance-hiding Schemes for Private Distributed LearningYangsibo Huang, Zhao Song, Kai Li, Sanjeev AroraICML 2020 · 178 citations
Related papers
- Towards Understanding the Data Dependency of Mixup-style TrainingMuthu Chidambaram, Xiang Wang, Yuzheng Hu, Chenwei Wu et al.ICLR 2022 · 25 citations
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani et al.ICLR 2021 · 294 citations
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 11 citations
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
