Global Mixup: Eliminating Ambiguity with Clustering
Xiangjin Xie, Yangning Li, Wang Chen, Kai Ouyang, Zuotong Xie, Haitao Zheng
Abstract
Data augmentation with Mixup has been proven an effective method to regularize the current deep neural networks. Mixup generates virtual samples and corresponding labels at once through linear interpolation. However, this one-stage generation paradigm and the use of linear interpolation have the following two defects: (1) The label of the generated sample is directly combined from the labels of the original sample pairs without reasonable judgment, which makes the labels likely to be ambiguous. ( 2 ) linear combination significantly limits the sampling space for generating samples. To tackle these problems, we propose a novel and effective augmentation method based on global clustering relationships named Global Mixup. Specifically, we transform the previous one-stage augmentation process into two-stage, decoupling the process of generating virtual samples from the labeling. And for the labels of the generated samples, relabeling is performed based on clustering by calculating the global relationships of the generated samples. In addition, we are no longer limited to linear relationships but generate more reliable virtual samples in a larger sampling space. Extensive experiments for CNN, LSTM, and BERT on five tasks show that Global Mixup significantly outperforms previous stateof-the-art baselines. Further experiments also demonstrate the advantage of Global Mixup in low-resource scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad0662d3-abca-489b-b0d8-206ceeda2063Builds on1
Related papers
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 124 citations
- Adversarial Mixing Policy for Relaxing Locally Linear Constraints in MixupGuang Liu, Yuzhao Mao, Hailong Huang, Weiguo Gao et al.EMNLP 2021 · 3 citations
- Embedding Space Interpolation Beyond Mini-Batch, Beyond Pairs and Beyond ExamplesShashanka Venkataramanan, Ewa Kijak, Laurent Amsaleg, Yannis AvrithisNeurIPS 2023 · 7 citations
- GenLabel: Mixup Relabeling using Generative ModelsJy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon et al.ICML 2022 · 15 citations
- KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLPYufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng et al.ICLR 2023 · 2 citations
