Do Not Have Enough Data? Deep Learning to the Rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, Naama Zwerdling
Abstract
Based on recent advances in natural language modeling and those in text generation capabilities, we propose a novel data augmentation method for text classification tasks. We use a powerful pre-trained neural network model to artificially synthesize new labeled data for supervised learning. We mainly focus on cases with scarce labeled data. Our method, referred to as language-model-based data augmentation (LAM-BADA), involves fine-tuning a state-of-the-art language generator to a specific task through an initial training phase on the existing (usually small) labeled data. Using the fine-tuned model and given a class label, new sentences for the class are generated. Our process then filters these new sentences by using a classifier trained on the original data. In a series of experiments, we show that LAMBADA improves classifiers performance on a variety of datasets. Moreover, LAM-BADA significantly improves upon the state-of-the-art techniques for data augmentation, specifically those applicable to text classification tasks with little data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c97562a3-5c40-4f45-9734-1629bfba2ea7Cited by top-tier papers30
- SeqMix: Augmenting Active Sequence Labeling via Sequence MixupRongzhi Zhang, Yue Yu, Chao ZhangEMNLP 2020 · 65 citations
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionMartin Josifoski, Marija Sakota, Maxime Peyrard, Robert WestEMNLP 2023 · 43 citations
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training DataSilei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. LamEMNLP 2020 · 33 citations
- Cluster & Tune: Boost Cold Start Performance in Text ClassificationEyal Shnarch, Ariel Gera, Alon Halfon, Lena Dankin et al.ACL 2022 · 27 citations
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian et al.EMNLP 2023 · 23 citations
Related papers
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian et al.ACL 2022 · 91 citations
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu et al.ACL 2022
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai et al.EMNLP 2020 · 132 citations
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma et al.EMNLP 2020 · 62 citations
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma et al.ICLR 2025
