Global Mixup: Eliminating Ambiguity with Clustering
Xiangjin Xie, Yangning Li, Wang Chen, Kai Ouyang, Zuotong Xie, Haitao Zheng
摘要
Data augmentation with Mixup has been proven an effective method to regularize the current deep neural networks. Mixup generates virtual samples and corresponding labels at once through linear interpolation. However, this one-stage generation paradigm and the use of linear interpolation have the following two defects: (1) The label of the generated sample is directly combined from the labels of the original sample pairs without reasonable judgment, which makes the labels likely to be ambiguous. ( 2 ) linear combination significantly limits the sampling space for generating samples. To tackle these problems, we propose a novel and effective augmentation method based on global clustering relationships named Global Mixup. Specifically, we transform the previous one-stage augmentation process into two-stage, decoupling the process of generating virtual samples from the labeling. And for the labels of the generated samples, relabeling is performed based on clustering by calculating the global relationships of the generated samples. In addition, we are no longer limited to linear relationships but generate more reliable virtual samples in a larger sampling space. Extensive experiments for CNN, LSTM, and BERT on five tasks show that Global Mixup significantly outperforms previous stateof-the-art baselines. Further experiments also demonstrate the advantage of Global Mixup in low-resource scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 被引用 124 次
- Adversarial Mixing Policy for Relaxing Locally Linear Constraints in MixupGuang Liu, Yuzhao Mao, Hailong Huang, Weiguo Gao 等EMNLP 2021 · 被引用 3 次
- Embedding Space Interpolation Beyond Mini-Batch, Beyond Pairs and Beyond ExamplesShashanka Venkataramanan, Ewa Kijak, Laurent Amsaleg, Yannis AvrithisNeurIPS 2023 · 被引用 7 次
- GenLabel: Mixup Relabeling using Generative ModelsJy-yong Sohn, Liang Shang, Hongxu Chen, Jaekyun Moon 等ICML 2022 · 被引用 15 次
- KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLPYufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng 等ICLR 2023 · 被引用 2 次
