MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NER
Ran Zhou, Xin Li, Ruidan He, Lidong Bing, Erik Cambria, Luo Si, Chunyan Miao
摘要
Data augmentation is an effective solution to data scarcity in low-resource scenarios. However, when applied to token-level tasks such as NER, data augmentation methods often suffer from token-label misalignment, which leads to unsatsifactory performance. In this work, we propose Masked Entity Language Modeling (MELM) as a novel data augmentation framework for low-resource NER. To alleviate the token-label misalignment issue, we explicitly inject NER labels into sentence context, and thus the fine-tuned MELM is able to predict masked entity tokens by explicitly conditioning on their labels. Thereby, MELM generates high-quality augmented data with novel entities, which provides rich entity regularity knowledge and boosts NER performance. When training data from multiple languages are available, we also integrate MELM with code-mixing for further improvement. We demonstrate the effectiveness of MELM on monolingual, cross-lingual and multilingual NER across various low-resource levels. Experimental results show that our MELM presents substantial improvement over the baseline methods. 1 * Ran Zhou is under the Joint Ph.D. Program between Alibaba and Nanyang Technological University.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng 等EMNLP 2022 · 被引用 76 次
- PromptNER: Prompt Locating and Typing for Named Entity RecognitionYongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang 等ACL 2023 · 被引用 46 次
- SciMON: Scientific Inspiration Machines Optimized for NoveltyQingyun Wang, Doug Downey, Heng Ji, Tom HopeACL 2024 · 被引用 22 次
- Improving Self-training for Cross-lingual Named Entity Recognition with Contrastive and Prototype LearningRan Zhou, Xin Li, Lidong Bing, Erik Cambria 等ACL 2023 · 被引用 19 次
- ConNER: Consistency Training for Cross-lingual Named Entity RecognitionRan Zhou, Xin Li, Lidong Bing, Erik Cambria 等EMNLP 2022 · 被引用 16 次
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DAGA: Data Augmentation with a Generation Approach forLow-resource Tagging TasksBosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai 等EMNLP 2020 · 被引用 132 次
- Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence GenerationKun Li, Chengbo Chen, Xiaojun Quan, Qing Ling 等ACL 2020 · 被引用 101 次
- Cross-lingual Aspect-based Sentiment Analysis with Aspect Term Code-SwitchingWenxuan Zhang, Ruidan He, Haiyun Peng, Lidong Bing 等EMNLP 2021 · 被引用 41 次
- A Rigorous Study on Named Entity Recognition: Can Fine-tuning Pretrained Model Lead to the Promised Land?Hongyu Lin, Yaojie Lu, Jialong Tang, Xianpei Han 等EMNLP 2020 · 被引用 41 次
相关 Paper
- ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NERSreyan Ghosh, Utkarsh Tyagi, Manan Suri, Sonal Kumar 等ACL 2023 · 被引用 9 次
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 被引用 7 次
- Generative Multimodal Data Augmentation for Low-Resource Multimodal Named Entity RecognitionZiyan Li, Jianfei Yu, Jia Yang, Wenya Wang 等ACM MM 2024 · 被引用 13 次
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 被引用 39 次
- MulDA: A Multilingual Data Augmentation Framework for Low-Resource Cross-Lingual NERLinlin Liu, Bosheng Ding, Lidong Bing, Shafiq R. Joty 等ACL 2021
