DropMix: A Textual Data Augmentation Combining Dropout with Mixup
Fanshuang Kong, Richong Zhang, Xiaohui Guo, Samuel Mensah, Yongyi Mao
摘要
Overfitting is a notorious problem when there is insufficient data to train deep neural networks in machine learning tasks. Data augmentation regularization methods such as Dropout, Mixup, and their enhanced variants are effective and prevalent, and achieve promising performance to overcome overfitting. However, in text learning, most of the existing regularization approaches merely adopt ideas from computer vision without considering the importance of dimensionality in natural language processing. In this paper, we argue that the property is essential to overcome overfitting in text learning. Accordingly, we present a saliency map informed textual data augmentation and regularization framework, which combines Dropout and Mixup, namely DropMix, to mitigate the overfitting problem in text learning. In addition, we design a procedure that drops and patches fine grained shapes of the saliency map under the DropMix framework to enhance regularization. Empirical studies confirm the effectiveness of the proposed approach on 12 text classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- iGraphMix: Input Graph Mixup Method for Node ClassificationJongwon Jeong, Hoyeop Lee, Hyui Geon Yoon, Beomyoung Lee 等ICLR 2024 · 被引用 10 次
- CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsJaehyung Seo, Hyeonseok Moon, Jaewook Lee, Sugyeong Eo 等EMNLP 2023 · 被引用 1 次
- ALVIN: Active Learning Via INterpolationMichalis Korakakis, Andreas Vlachos, Adrian WellerEMNLP 2024
- Mitigating Shortcut Learning with InterpoLated LearningMichalis Korakakis, Andreas Vlachos, Adrian WellerACL 2025
它引用的顶会 Paper5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 被引用 457 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- SaliencyMix: A Saliency Guided Data Augmentation Strategy for Better RegularizationA. F. M. Shahab Uddin, Mst. Sirazam Monira, Wheemyung Shin, TaeChoong Chung 等ICLR 2021 · 被引用 271 次
- SeqMix: Augmenting Active Sequence Labeling via Sequence MixupRongzhi Zhang, Yue Yu, Chao ZhangEMNLP 2020 · 被引用 65 次
相关 Paper
- Catch-Up Mix: Catch-Up Class for Struggling Filters in CNNMinsoo Kang, Minkoo Kang, Suhyun KimAAAI 2024 · 被引用 7 次
- Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksHaoqi Zheng, Qihuang Zhong, Liang Ding, Zhiliang Tian 等EMNLP 2023 · 被引用 4 次
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsCheolhyoung Lee, Kyunghyun Cho, Wanmo KangICLR 2020 · 被引用 233 次
- Regularization Strategy for Point Cloud via Rigidly Mixed SampleDogyoon Lee, Jaeha Lee, Junhyeop Lee, Hyeongmin Lee 等CVPR 2021
- TLM: Token-Level Masking for TransformersYangjun Wu, Kebin Fang, Dongxiang Zhang, Han Wang 等EMNLP 2023 · 被引用 2 次
