Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled Data
David Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. Wallace
摘要
Unsupervised Data Augmentation (UDA) is a semi-supervised technique that applies a consistency loss to penalize differences between a model's predictions on (a) observed (unlabeled) examples; and (b) corresponding 'noised' examples produced via data augmentation. While UDA has gained popularity for text classification, open questions linger over which design decisions are necessary and over how to extend the method to sequence labeling tasks. In this paper, we re-examine UDA and demonstrate its efficacy on several sequential tasks. Our main contribution is an empirical study of UDA to establish which components of the algorithm confer benefits in NLP. Notably, although prior work has emphasized the use of clever augmentation techniques including back-translation, we find that enforcing consistency between predictions assigned to observed and randomly substituted words often yields comparable (or greater) benefits compared to these complex perturbation models. Furthermore, we find that applying its consistency loss affords meaningful gains without any unlabeled data at all, i.e., in a standard supervised setting. In short: UDA need not be unsupervised, and does not require complex data augmentation to be effective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao 等EMNLP 2022 · 被引用 20 次
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria 等EMNLP 2022 · 被引用 16 次
- CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme IngredientsJaehyung Seo, Hyeonseok Moon, Jaewook Lee, Sugyeong Eo 等EMNLP 2023 · 被引用 1 次
- Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency LearningHao Wang, Xiahua Chen, Rui Wang, Chenhui ChuEMNLP 2023
它引用的顶会 Paper1
相关 Paper
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 被引用 294 次
- Local Additivity Based Data Augmentation for Semi-supervised NERJiaao Chen, Zhenghui Wang, Ran Tian, Zichao Yang 等EMNLP 2020 · 被引用 45 次
- Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalDevang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva ReddyEMNLP 2021 · 被引用 11 次
- Consistency Regularization for Cross-Lingual Fine-TuningBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang 等ACL 2021
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han 等CVPR 2022 · 被引用 20 次
