The Benefits of Mixup for Feature Learning
Difan Zou, Yuan Cao, Yuanzhi Li, Quanquan Gu
摘要
Mixup, a simple data augmentation method that randomly mixes two data points via linear interpolation, has been extensively applied in various deep learning applications to gain better generalization. However, the theoretical underpinnings of its efficacy are not yet fully understood. In this paper, we aim to seek a fundamental understanding of the benefits of Mixup. We first show that Mixup using different linear interpolation parameters for features and labels can still achieve similar performance to the standard Mixup. This indicates that the intuitive linearity explanation in Zhang et al. ( 2018 ) may not fully explain the success of Mixup. Then we perform a theoretical study of Mixup from the feature learning perspective. We consider a feature-noise data model and show that Mixup training can effectively learn the rare features (appearing in a small fraction of data) from its mixture with the common features (appearing in a large fraction of data). In contrast, standard training can only learn the common features but fails to learn the rare features, thus suffering from bad generalization performance. Moreover, our theoretical analysis also shows that the benefits of Mixup for feature learning are mostly gained in the early training phase, based on which we propose to apply early stopping in Mixup. Experimental results verify our theoretical findings and demonstrate the effectiveness of the early-stopped Mixup training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Toward Understanding Generative Data AugmentationChenyu Zheng, Guoqiang Wu, Chongxuan LiNeurIPS 2023 · 被引用 51 次
- Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and MethodBikang Pan, Wei Huang, Ye ShiNeurIPS 2024 · 被引用 28 次
- Understanding Convergence and Generalization in Federated Learning through Feature Learning TheoryWei Huang, Ye Shi, Zhongyi Cai, Taiji SuzukiICLR 2024 · 被引用 17 次
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 被引用 15 次
- Provable Benefit of Cutout and CutMix for Feature LearningJunsoo Oh, Chulhee YunNeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper19
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 被引用 254 次
- G-Mixup: Graph Data Augmentation for Graph ClassificationXiaotian Han, Zhimeng Jiang, Ninghao Liu, Xia HuICML 2022 · 被引用 251 次
相关 Paper
- Over-Training with Mixup May Hurt GeneralizationZixuan Liu, Ziqiao Wang, Hongyu Guo, Yongyi MaoICLR 2023 · 被引用 2 次
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 被引用 13 次
- Harnessing Hard Mixed Samples with Decoupled RegularizerZicheng Liu, Siyuan Li, Ge Wang, Lirong Wu 等NeurIPS 2023 · 被引用 28 次
- Towards Understanding the Data Dependency of Mixup-style TrainingMuthu Chidambaram, Xiang Wang, Yuzheng Hu, Chenwei Wu 等ICLR 2022 · 被引用 25 次
- RC-Mixup: A Data Augmentation Strategy against Noisy Data for Regression TasksSeonghyeon Hwang, Minsu Kim, Steven Euijong WhangKDD 2024 · 被引用 4 次
