Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled Data
David Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. Wallace
Abstract
Unsupervised Data Augmentation (UDA) is a semi-supervised technique that applies a consistency loss to penalize differences between a model's predictions on (a) observed (unlabeled) examples; and (b) corresponding 'noised' examples produced via data augmentation. While UDA has gained popularity for text classification, open questions linger over which design decisions are necessary and over how to extend the method to sequence labeling tasks. In this paper, we re-examine UDA and demonstrate its efficacy on several sequential tasks. Our main contribution is an empirical study of UDA to establish which components of the algorithm confer benefits in NLP. Notably, although prior work has emphasized the use of clever augmentation techniques including back-translation, we find that enforcing consistency between predictions assigned to observed and randomly substituted words often yields comparable (or greater) benefits compared to these complex perturbation models. Furthermore, we find that applying its consistency loss affords meaningful gains without any unlabeled data at all, i.e., in a standard supervised setting. In short: UDA need not be unsupervised, and does not require complex data augmentation to be effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19323881-e627-4fdb-a7ee-aba866ea3c51Cited by top-tier papers4
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria et al.EMNLP 2022 · 16 citations
- CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsJaehyung Seo, Hyeonseok Moon, Jaewook Lee, Sugyeong Eo et al.EMNLP 2023 · 1 citation
- Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency LearningHao Wang, Xiahua Chen, Rui Wang, Chenhui ChuEMNLP 2023
Builds on1
Related papers
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- Local Additivity Based Data Augmentation for Semi-supervised NERJiaao Chen, Zhenghui Wang, Ran Tian, Zichao Yang et al.EMNLP 2020 · 45 citations
- Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalDevang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva ReddyEMNLP 2021 · 11 citations
- Consistency Regularization for Cross-Lingual Fine-TuningBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang et al.ACL 2021
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
