Over-Training with Mixup May Hurt Generalization
Zixuan Liu, Ziqiao Wang, Hongyu Guo, Yongyi Mao
摘要
Mixup, which creates synthetic training instances by linearly interpolating random sample pairs, is a simple and yet effective regularization technique to boost the performance of deep models trained with SGD. In this work, we report a previously unobserved phenomenon in Mixup training: on a number of standard datasets, the performance of Mixup-trained models starts to decay after training for a large number of epochs, giving rise to a U-shaped generalization curve. This behavior is further aggravated when the size of original dataset is reduced. To help understand such a behavior of Mixup, we show theoretically that Mixup training may introduce undesired data-dependent label noises to the synthesized data. Via analyzing a least-square regression problem with a random feature model, we explain why noisy labels may cause the U-shaped curve to occur: Mixup improves generalization through fitting the clean patterns at the early training stage, but as training progresses, Mixup becomes over-fitting to the noise in the synthetic data. Extensive experiments are performed on a variety of benchmark datasets, validating this explanation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion PriorsPaul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin 等NeurIPS 2023 · 被引用 282 次
- Semi-Supervised Graph Imbalanced RegressionGang Liu, Tong Zhao, Eric Inae, Tengfei Luo 等KDD 2023 · 被引用 20 次
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of MixupDamien Teney, Jindong Wang, Ehsan AbbasnejadICML 2024 · 被引用 9 次
- Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and BeyondThanh Vu, Baochen Sun, Bodi Yuan, Alex Ngai 等AAAI 2024 · 被引用 8 次
- Denoising Mixup for RegressionZhengzhang Hou, Zhanshan Li, Yanbo Liu, Geoff Nitschke 等AAAI 2026
它引用的顶会 Paper16
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
相关 Paper
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 被引用 36 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- GenLabel: Mixup Relabeling using Generative ModelsJy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon 等ICML 2022 · 被引用 15 次
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 被引用 9 次
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
