Provable Benefit of Mixup for Finding Optimal Decision Boundaries
Junsoo Oh, Chulhee Yun
摘要
We investigate how pair-wise data augmentation techniques like Mixup affect the sample complexity of finding optimal decision boundaries in a binary linear classification problem. For a family of data distributions with a separability constant , we analyze how well the optimal classifier in terms of training loss aligns with the optimal one in test accuracy (i.e., Bayes optimal classifier). For vanilla training without augmentation, we uncover an interesting phenomenon named the curse of separability. As we increase to make the data distribution more separable, the sample complexity of vanilla training increases exponentially in ; perhaps surprisingly, the task of finding optimal decision boundaries becomes harder for more separable distributions. For Mixup training, we show that Mixup mitigates this problem by significantly reducing the sample complexity. To this end, we develop new concentration results applicable to pair-wise augmented data points constructed from independent data, by carefully dealing with dependencies between overlapping pairs. Lastly, we study other masking-based Mixup-style techniques and show that they can distort the training loss and make its minimizer converge to a suboptimal classifier in terms of test accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On Unsupervised Domain Adaptation: Pseudo Label Guided Mixup for Adversarial Prompt TuningFanshuang Kong, Richong Zhang, Ziqiao Wang, Yongyi MaoAAAI 2024 · 被引用 12 次
- Provable Benefit of Cutout and CutMix for Feature LearningJunsoo Oh, Chulhee YunNeurIPS 2024 · 被引用 11 次
- SaVe-TAG: LLM-based Interpolation for Long-Tailed Text-Attributed GraphsLeyao Wang, Yu Wang, Bo Ni, Yuying Zhao 等KDD 2026 · 被引用 1 次
- For Better or For Worse? Learning Minimum Variance Features With Label AugmentationMuthu Chidambaram, Rong GeICLR 2025
- Tailoring Mixup to Data for CalibrationQuentin Bouniot, Pavlo Mozharovskyi, Florence d'Alché-BucICLR 2025
它引用的顶会 Paper9
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel 等NeurIPS 2020 · 被引用 805 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville 等NeurIPS 2021 · 被引用 378 次
- Co-Mixup: Saliency Guided Joint Mixup with Supermodular DiversityJang-Hyun Kim, Wonho Choo, Hosan Jeong, Hyun Oh SongICLR 2021 · 被引用 207 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
相关 Paper
- Towards Understanding the Data Dependency of Mixup-style TrainingMuthu Chidambaram, Xiang Wang, Yuzheng Hu, Chenwei Wu 等ICLR 2022 · 被引用 25 次
- Harnessing Hard Mixed Samples with Decoupled RegularizerZicheng Liu, Siyuan Li, Ge Wang, Lirong Wu 等NeurIPS 2023 · 被引用 28 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- Boundary thickness and robustness in learning modelsYaoqing Yang, Rajiv Khanna, Yaodong Yu, Amir Gholami 等NeurIPS 2020 · 被引用 53 次
- Selective Mixup Helps with Distribution Shifts, But Not (Only) because of MixupDamien Teney, Jindong Wang, Ehsan AbbasnejadICML 2024 · 被引用 9 次
