Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition
Shuguang Chen, Leonardo Neves, Thamar Solorio
Abstract
In this work, we take the named entity recognition task in the English language as a case study and explore style transfer as a data augmentation method to increase the size and diversity of training data in low-resource scenarios. We propose a new method to effectively transform the text from a high-resource domain to a low-resource domain by changing its style-related attributes to generate synthetic data for training. Moreover, we design a constrained decoding algorithm along with a set of key ingredients for data selection to guarantee the generation of valid and coherent data. Experiments and analysis on five different domain pairs under different data regimes demonstrate that our approach can significantly improve results compared to current state-of-the-art data augmentation methods. Our approach is a practical solution to data scarcity, and we expect it to be applicable to other NLP tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a78646b-8822-4d15-898b-b442b8ab67d7Cited by top-tier papers5
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 7 citations
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru et al.ACL 2024 · 3 citations
- Exogenous and Endogenous Data Augmentation for Low-Resource Complex Named Entity RecognitionXinghua Zhang, Gaode Chen, Shiyao Cui, Jiawei Sheng et al.SIGIR 2024 · 3 citations
- Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?Wataru Hashimoto, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 1 citation
- RSDA: Restoring Stale Data Affinity via Dynamic Renovation Strategy for Mitigating Data ScarcityYidan Liang, Jia Zhu, Weijie Shi, Hanghui Guo et al.ACL 2026
Builds on26
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
Related papers
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 39 citations
- MulDA: A Multilingual Data Augmentation Framework for Low-Resource Cross-Lingual NERLinlin Liu, Bosheng Ding, Lidong Bing, Shafiq R. Joty et al.ACL 2021
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing et al.ACL 2022 · 114 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai et al.EMNLP 2020 · 132 citations
