For Better or For Worse? Learning Minimum Variance Features With Label Augmentation
Muthu Chidambaram, Rong Ge
摘要
Data augmentation has been pivotal in successfully training deep learning models on classification tasks over the past decade. An important subclass of data augmentation techniques - which includes both label smoothing and Mixup - involves modifying not only the input data but also the input label during model training. In this work, we analyze the role played by the label augmentation aspect of such methods. We first prove that linear models on binary classification data trained with label augmentation learn only the minimum variance features in the data, while standard training (which includes weight decay) can learn higher variance features. We then use our techniques to show that even for nonlinear models and general data distributions, the label smoothing and Mixup losses are lower bounded by a function of the model output variance. Lastly, we demonstrate empirically that this aspect of label smoothing and Mixup can be a positive and a negative. On the one hand, we show that the strong performance of label smoothing and Mixup on image classification benchmarks is correlated with learning low variance hidden representations. On the other hand, we show that Mixup and label smoothing can be more susceptible to low variance spurious correlations in the training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper24
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
相关 Paper
- Provably Learning Diverse Features in Multi-View Data with Midpoint MixupMuthu Chidambaram, Xiang Wang, Chenwei Wu, Rong GeICML 2023 · 被引用 13 次
- When Does Data Augmentation Help With Membership Inference Attacks?Yigitcan Kaya, Tudor DumitrasICML 2021 · 被引用 82 次
- Over-Training with Mixup May Hurt GeneralizationZixuan Liu, Ziqiao Wang, Hongyu Guo, Yongyi MaoICLR 2023 · 被引用 2 次
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani 等ICLR 2021 · 被引用 294 次
- Towards Understanding the Data Dependency of Mixup-style TrainingMuthu Chidambaram, Xiang Wang, Yuzheng Hu, Chenwei Wu 等ICLR 2022 · 被引用 25 次
