Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training
Yu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang, Yu Zhang, Heng Ji, Jiawei Han
摘要
We study the problem of training named entity recognition (NER) models using only distantly-labeled data, which can be automatically obtained by matching entity mentions in the raw text with entity types in a knowledge base. The biggest challenge of distantlysupervised NER is that the distant supervision may induce incomplete and noisy labels, rendering the straightforward application of supervised learning ineffective. In this paper, we propose (1) a noise-robust learning scheme comprised of a new loss function and a noisy label removal step, for training NER models on distantly-labeled data, and (2) a self-training method that uses contextualized augmentations created by pre-trained language models to improve the generalization ability of the NER model. On three benchmark datasets, our method achieves superior performance, outperforming existing distantlysupervised NER models by significant margins 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NERDong-Ho Lee, Akshen Kadakia, Kangmin Tan, Mahak Agarwal 等ACL 2022 · 被引用 96 次
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang 等ICML 2023 · 被引用 64 次
- Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance GenerationJiaxin Huang, Yu Meng, Jiawei HanKDD 2022 · 被引用 17 次
- PIEClass: Weakly-Supervised Text Classification with Prompting and Noise-Robust Iterative Ensemble TrainingYunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang 等EMNLP 2023 · 被引用 16 次
它引用的顶会 Paper10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- SELF: Learning to Filter Noisy Labels with Self-EnsemblingDuc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen 等ICLR 2020 · 被引用 354 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等NeurIPS 2021 · 被引用 231 次
相关 Paper
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Coarse-to-Fine Pre-training for Named Entity RecognitionMengge Xue, Bowen Yu, Zhenyu Zhang, Tingwen Liu 等EMNLP 2020 · 被引用 49 次
- MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity RecognitionShuhui Wu, Yongliang Shen, Zeqi Tan, Wenqi Ren 等EMNLP 2023 · 被引用 4 次
- Empirical Analysis of Unlabeled Entity Problem in Named Entity RecognitionYangming Li, Lemao Liu, Shuming ShiICLR 2021 · 被引用 72 次
- Denoising Distantly Supervised Named Entity Recognition via a Hypergeometric Probabilistic ModelWenkai Zhang, Hongyu Lin, Xianpei Han, Le Sun 等AAAI 2021 · 被引用 13 次
