Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled Data
Haoming Jiang, Danqing Zhang, Tianyu Cao, Bing Yin, Tuo Zhao
摘要
Weak supervision has shown promising results in many natural language processing tasks, such as Named Entity Recognition (NER). Existing work mainly focuses on learning deep NER models only with weak supervision, i.e., without any human annotation, and shows that by merely using weakly labeled data, one can achieve good performance, though still underperforms fully supervised NER with manually/strongly labeled data. In this paper, we consider a more practical scenario, where we have both a small amount of strongly labeled data and a large amount of weakly labeled data. Unfortunately, we observe that weakly labeled data does not necessarily improve, or even deteriorate the model performance (due to the extensive noise in the weak labels) when we train deep NER models over a simple or weighted combination of the strongly labeled and weakly labeled data. To address this issue, we propose a new multi-stage computational framework -NEEDLE with three essential ingredients: (1) weak label completion, (2) noise-aware loss function, and (3) final finetuning over the strongly labeled data. Through experiments on E-commerce query NER and Biomedical NER, we demonstrate that NEE-DLE can effectively suppress the noise of the weak labels and outperforms existing methods. In particular, we achieve new SOTA F1-scores on 3 Biomedical NER datasets: BC5CDRchem 93.74, BC5CDR-disease 90.69, NCBIdisease 92.28.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu 等AAAI 2022 · 被引用 36 次
- Debiased and Denoised Entity Recognition from Distant SupervisionHaobo Wang, Yiwen Dong, Ruixuan Xiao, Fei Huang 等NeurIPS 2023 · 被引用 5 次
- Enhancing Low-resource Fine-grained Named Entity Recognition by Leveraging Coarse-grained DatasetsSu Ah Lee, Seokjin Oh, Woohwan JungEMNLP 2023 · 被引用 4 次
- Refining and Reusing Annotation Guidelines for LLM AnnotationKon Woo Kim, Jin-Dong Kim, Akiko AizawaACL 2026
- Few-shot Named Entity Recognition with Self-describing NetworksJiawei Chen, Qing Liu, Hongyu Lin, Xianpei Han 等ACL 2022
它引用的顶会 Paper4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingOuyu Lan, Xiao Huang, Bill Yuchen Lin, He Jiang 等ACL 2020 · 被引用 33 次
相关 Paper
- Named Entity Recognition without Labelled Data: A Weak Supervision ApproachPierre Lison, Jeremy Barnes, Aliaksandr Hubin, Samia TouilebACL 2020 · 被引用 12 次
- BERTifying the Hidden Markov Model for Multi-Source Weakly Supervised Named Entity RecognitionYinghao Li, Pranav Shetty, Lucas Liu, Chao Zhang 等ACL 2021
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang 等EMNLP 2021 · 被引用 50 次
- Named Entity Recognition Only from Word EmbeddingsYing Luo, Hai Zhao, Junlang ZhanEMNLP 2020 · 被引用 22 次
- Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsJian Liu, Weichang Liu, Yufeng Chen, Jinan Xu 等EMNLP 2023 · 被引用 3 次
