Over-Training with Mixup May Hurt Generalization
Zixuan Liu, Ziqiao Wang, Hongyu Guo, Yongyi Mao
Abstract
Mixup, which creates synthetic training instances by linearly interpolating random sample pairs, is a simple and yet effective regularization technique to boost the performance of deep models trained with SGD. In this work, we report a previously unobserved phenomenon in Mixup training: on a number of standard datasets, the performance of Mixup-trained models starts to decay after training for a large number of epochs, giving rise to a U-shaped generalization curve. This behavior is further aggravated when the size of original dataset is reduced. To help understand such a behavior of Mixup, we show theoretically that Mixup training may introduce undesired data-dependent label noises to the synthesized data. Via analyzing a least-square regression problem with a random feature model, we explain why noisy labels may cause the U-shaped curve to occur: Mixup improves generalization through fitting the clean patterns at the early training stage, but as training progresses, Mixup becomes over-fitting to the noise in the synthetic data. Extensive experiments are performed on a variety of benchmark datasets, validating this explanation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bafbbf84-e2f4-4752-91ac-8f7c94f3815dCited by top-tier papers7
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin et al.NeurIPS 2023 · 282 citations
- Semi-Supervised Graph Imbalanced RegressionGang Liu, Tong Zhao, Eric Inae, Tengfei Luo et al.KDD 2023 · 20 citations
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of MixupDamien Teney, Jindong Wang, Ehsan AbbasnejadICML 2024 · 9 citations
- Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and BeyondThanh Vu, Baochen Sun, Bodi Yuan, Alex Ngai et al.AAAI 2024 · 8 citations
- Denoising Mixup for RegressionZhengzhang Hou, Zhanshan Li, Yanbo Liu, Geoff Nitschke et al.AAAI 2026
Builds on16
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 457 citations
Related papers
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 36 citations
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani et al.ICLR 2021 · 294 citations
- GenLabel: Mixup Relabeling using Generative ModelsJy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon et al.ICML 2022 · 15 citations
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 9 citations
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
