Automatic Creation of Named Entity Recognition Datasets by Querying Phrase Representations
Hyunjae Kim, Jaehyo Yoo, Seunghyun Yoon, Jaewoo Kang
摘要
Most weakly supervised named entity recognition (NER) models rely on domain-specific dictionaries provided by experts. This approach is infeasible in many domains where dictionaries do not exist. While a phrase retrieval model was used to construct pseudo-dictionaries with entities retrieved from Wikipedia automatically in a recent study, these dictionaries often have limited coverage because the retriever is likely to retrieve popular entities rather than rare ones. In this study, we present a novel framework, HighGEN, that generates NER datasets with high-coverage pseudo-dictionaries. Specifically, we create entity-rich dictionaries with a novel search method, called phrase embedding search, which encourages the retriever to search a space densely populated with various entities. In addition, we use a new verification process based on the embedding distance between candidate entity mentions and entity types to reduce the false-positive noise in weak labels generated by high-coverage dictionaries. We demonstrate that HighGEN outperforms the previous best model by an average F1 score of 4.7 across five NER benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Simple and Effective Few-Shot Named Entity Recognition with Structured Nearest Neighbor LearningYi Yang, Arzoo KatiyarEMNLP 2020 · 被引用 198 次
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Weakly Supervised Sequence Tagging from Noisy RulesEsteban Safranchik, Shiying Luo, Stephen H. BachAAAI 2020 · 被引用 90 次
相关 Paper
- Simple Questions Generate Named Entity Recognition DatasetsHyunjae Kim, Jaehyo Yoo, Seunghyun Yoon, Jinhyuk Lee 等EMNLP 2022 · 被引用 4 次
- HAMNER: Headword Amplified Multi-Span Distantly Supervised Method for Domain Specific Named Entity RecognitionShifeng Liu, Yifang Sun, Bing Li, Wei Wang 等AAAI 2020 · 被引用 28 次
- A Rigorous Study on Named Entity Recognition: Can Fine-tuning Pretrained Model Lead to the Promised Land?Hongyu Lin, Yaojie Lu, Jialong Tang, Xianpei Han 等EMNLP 2020 · 被引用 41 次
- Coarse-to-Fine Pre-training for Named Entity RecognitionMengge Xue, Bowen Yu, Zhenyu Zhang, Tingwen Liu 等EMNLP 2020 · 被引用 49 次
- Named Entity Recognition Only from Word EmbeddingsYing Luo, Hai Zhao, Junlang ZhanEMNLP 2020 · 被引用 22 次
