Do Not Have Enough Data? Deep Learning to the Rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, Naama Zwerdling
摘要
Based on recent advances in natural language modeling and those in text generation capabilities, we propose a novel data augmentation method for text classification tasks. We use a powerful pre-trained neural network model to artificially synthesize new labeled data for supervised learning. We mainly focus on cases with scarce labeled data. Our method, referred to as language-model-based data augmentation (LAM-BADA), involves fine-tuning a state-of-the-art language generator to a specific task through an initial training phase on the existing (usually small) labeled data. Using the fine-tuned model and given a class label, new sentences for the class are generated. Our process then filters these new sentences by using a classifier trained on the original data. In a series of experiments, we show that LAMBADA improves classifiers performance on a variety of datasets. Moreover, LAM-BADA significantly improves upon the state-of-the-art techniques for data augmentation, specifically those applicable to text classification tasks with little data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- SeqMix: Augmenting Active Sequence Labeling via Sequence MixupRongzhi Zhang, Yue Yu, Chao ZhangEMNLP 2020 · 被引用 65 次
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionMartin Josifoski, Marija Sakota, Maxime Peyrard, Robert WestEMNLP 2023 · 被引用 43 次
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training DataSilei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. LamEMNLP 2020 · 被引用 33 次
- Cluster & Tune: Boost Cold Start Performance in Text ClassificationEyal Shnarch, Ariel Gera, Alon Halfon, Lena Dankin 等ACL 2022 · 被引用 27 次
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian 等EMNLP 2023 · 被引用 23 次
相关 Paper
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian 等ACL 2022 · 被引用 91 次
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu 等ACL 2022
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai 等EMNLP 2020 · 被引用 132 次
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma 等EMNLP 2020 · 被引用 62 次
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
