Robust Named Entity Recognition with Truecasing Pretraining
Stephen Mayhew, Nitish Gupta, Dan Roth
摘要
Although modern named entity recognition (NER) systems show impressive performance on standard datasets, they perform poorly when presented with noisy data. In particular, capitalization is a strong signal for entities in many languages, and even state of the art models overfit to this feature, with drastically lower performance on uncapitalized text. In this work, we address the problem of robustness of NER systems in data with noisy or uncertain casing, using a pretraining objective that predicts casing in text, or a truecaser, leveraging unlabeled data. The pretrained truecaser is combined with a standard BiLSTM-CRF model for NER by appending output distributions to character embeddings. In experiments over several datasets of varying domain and casing quality, we show that our new model improves performance in uncased text, even adding value to uncased BERT embeddings. Our method achieves a new state of the art on the WNUT17 shared task dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Few-Shot Text Generation with Natural Language InstructionsTimo Schick, Hinrich SchützeEMNLP 2021 · 被引用 101 次
- Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang 等ACL 2021
相关 Paper
- BERTifying the Hidden Markov Model for Multi-Source Weakly Supervised Named Entity RecognitionYinghao Li, Pranav Shetty, Lucas Liu, Chao Zhang 等ACL 2021
- Empirical Analysis of Unlabeled Entity Problem in Named Entity RecognitionYangming Li, Lemao Liu, Shuming ShiICLR 2021 · 被引用 72 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Entity Enhanced BERT Pre-training for Chinese NERChen Jia, Yuefeng Shi, Qinrong Yang, Yue ZhangEMNLP 2020 · 被引用 59 次
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang 等EMNLP 2021 · 被引用 50 次
