FlipDA: Effective and Robust Data Augmentation for Few-Shot Learning
Jing Zhou, Yanan Zheng, Jie Tang, Li Jian, Zhilin Yang
Abstract
Most previous methods for text data augmentation are limited to simple tasks and weak baselines. We explore data augmentation on hard tasks (i.e., few-shot natural language understanding) and strong baselines (i.e., pretrained models with over one billion parameters). Under this setting, we reproduced a large number of previous augmentation methods and found that these methods bring marginal gains at best and sometimes degrade the performance much. To address this challenge, we propose a novel data augmentation method FlipDA that jointly uses a generative model and a classifier to generate label-flipped data. Central to the idea of FlipDA is the discovery that generating labelflipped data is more crucial to the performance than generating label-preserved data. Experiments show that FlipDA achieves a good tradeoff between effectiveness and robustness-it substantially improves many tasks while not negatively affecting the others. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbc7500b-0b8f-4f8a-8926-810f7e2b2f06Cited by top-tier papers18
- Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human InterventionsJohn Joon Young Chung, Ece Kamar, Saleema AmershiACL 2023 · 55 citations
- FewNLU: Benchmarking State-of-the-Art Methods for Few-Shot Natural Language UnderstandingYanan Zheng, Jing Zhou, Yujie Qian, Ming Ding et al.ACL 2022 · 31 citations
- Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text ClassificationPengyu Xu, Lin Xiao, Bing Liu, Sijin Lu et al.AAAI 2023 · 29 citations
- Domain Adaptive Few-Shot Open-Set LearningDebabrata Pal, Deeptej More, Sai Bhargav, Dipesh Tamboli et al.ICCV 2023 · 13 citations
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. et al.EMNLP 2023 · 10 citations
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
Related papers
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 7 citations
- ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text ClassificationHazel H. Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha et al.AAAI 2022 · 33 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai et al.EMNLP 2020 · 132 citations
