Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity Recognition
Hyunwoo Sohn, Baekkwan Park
Abstract
Recent advances in deep learning have brought remarkable performance improvements in named entity recognition (NER), specifically in token-level classification problems. However, deep learning models often require a large amount of annotated data to achieve satisfactory performance, and NER annotation is significantly time-consuming and labor-intensive due to the fine-grained labels. To address this issue, we propose a textual data augmentation method that can automatically generate informative synthetic samples, which contribute to the development of a robust classifier. The proposed method generates additional training data by estimating the optimal level of worst-case transformation of training data while preserving the original annotation, and includes them into training to construct a robust decision boundary. Extensive experiments conducted on two benchmark datasets in a low-resource environment reveal that the proposed method outperforms two baseline augmentation methods including human annotation, which is typically considered to provide a decent amount of performance boost. To elucidate the processes, we also present in-depth analyses of the generated samples and estimated model parameters.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai et al.EMNLP 2020 · 132 citations
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 7 citations
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 39 citations
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing et al.ACL 2022 · 114 citations
- Style Transfer as Data Augmentation: A Case Study on Named Entity RecognitionShuguang Chen, Leonardo Neves, Thamar SolorioEMNLP 2022 · 9 citations
