Towards Understanding the Data Dependency of Mixup-style Training
Muthu Chidambaram, Xiang Wang, Yuzheng Hu, Chenwei Wu, Rong Ge
摘要
In the Mixup training paradigm, a model is trained using convex combinations of data points and their associated labels. Despite seeing very few true data points during training, models trained using Mixup seem to still minimize the original empirical risk and exhibit better generalization and robustness on various tasks when compared to standard training. In this paper, we investigate how these benefits of Mixup training rely on properties of the data in the context of classification. For minimizing the original empirical risk, we compute a closed form for the Mixup-optimal classification, which allows us to construct a simple dataset on which minimizing the Mixup loss can provably lead to learning a classifier that does not minimize the empirical loss on the data. On the other hand, we also give sufficient conditions for Mixup training to also minimize the original empirical risk. For generalization, we characterize the margin of a Mixup classifier, and use this to understand why the decision boundary of a Mixup classifier can adapt better to the full structure of the training data when compared to standard training. In contrast, we also show that, for a large class of linear models and linearly separable datasets, Mixup training leads to learning the same classifier as standard training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- A Unified Analysis of Mixed Sample Data Augmentation: A Loss Function PerspectiveChanwoo Park, Sangdoo Yun, Sanghyuk ChunNeurIPS 2022 · 被引用 43 次
- The Benefits of Mixup for Feature LearningDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICML 2023 · 被引用 36 次
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 被引用 13 次
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 被引用 11 次
- Provable Benefit of Cutout and CutMix for Feature LearningJunsoo Oh, Chulhee YunNeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville 等NeurIPS 2021 · 被引用 378 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
相关 Paper
- GenLabel: Mixup Relabeling using Generative ModelsJy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon 等ICML 2022 · 被引用 15 次
- Boundary thickness and robustness in learning modelsYaoqing Yang, Rajiv Khanna, Yaodong Yu, Amir Gholami 等NeurIPS 2020 · 被引用 53 次
- On the Pitfall of Mixup for Uncertainty CalibrationDeng-Bao Wang, Lanqing Li, Peilin Zhao, Pheng-Ann Heng 等CVPR 2023
- Provable Benefit of Mixup for Finding Optimal Decision BoundariesJunsoo Oh, Chulhee YunICML 2023 · 被引用 6 次
- Over-Training with Mixup May Hurt GeneralizationZixuan Liu, Ziqiao Wang, Hongyu Guo, Yongyi MaoICLR 2023 · 被引用 2 次
