A Rigorous Study on Named Entity Recognition: Can Fine-tuning Pretrained Model Lead to the Promised Land?
Hongyu Lin, Yaojie Lu, Jialong Tang, Xianpei Han, Le Sun, Zhicheng Wei, Nicholas Jing Yuan
摘要
Fine-tuning pretrained model has achieved promising performance on standard NER benchmarks. Generally, these benchmarks are blessed with strong name regularity, high mention coverage and sufficient context diversity. Unfortunately, when scaling NER to open situations, these advantages may no longer exist. And therefore it raises a critical question of whether previous creditable approaches can still work well when facing these challenges. As there is no currently available dataset to investigate this problem, this paper proposes to conduct randomization test on standard benchmarks. Specifically, we erase name regularity, mention coverage and context diversity respectively from the benchmarks, in order to explore their impact on the generalization ability of models. To further verify our conclusions, we also construct a new open NER dataset that focuses on entity types with weaker name regularity and lower mention coverage to verify our conclusion. From both randomization test and empirical experiments, we draw the conclusions that 1) name regularity is critical for the models to generalize to unseen mentions; 2) high mention coverage may undermine the model generalization ability and 3) context patterns may not require enormous data to capture when using pretrained encoders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing 等ACL 2022 · 被引用 114 次
- LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using UncertaintyZhen Zhang, Yuhua Zhao, Hang Gao, Mengting HuWWW 2024 · 被引用 50 次
- MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou 等ACL 2022 · 被引用 35 次
- Denoising Distantly Supervised Named Entity Recognition via a Hypergeometric Probabilistic ModelWenkai Zhang, Hongyu Lin, Xianpei Han, Le Sun 等AAAI 2021 · 被引用 13 次
- Robustness of Demonstration-based Learning Under Limited Data ScenarioHongxin Zhang, Yanzhe Zhang, Ruiyi Zhang, Diyi YangEMNLP 2022 · 被引用 10 次
相关 Paper
- Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewRuotian Ma, Xiaolei Wang, Xin Zhou, Qi Zhang 等EMNLP 2023 · 被引用 3 次
- Automatic Creation of Named Entity Recognition Datasets by Querying Phrase RepresentationsHyunjae Kim, Jaehyo Yoo, Seunghyun Yoon, Jaewoo KangACL 2023 · 被引用 3 次
- Few-NERD: A Few-shot Named Entity Recognition DatasetNing Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang 等ACL 2021
- Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?Shuheng Liu, Alan RitterACL 2023 · 被引用 9 次
- CrossNER: Evaluating Cross-Domain Named Entity RecognitionZihan Liu, Yan Xu, Tiezheng Yu, Wenliang Dai 等AAAI 2021 · 被引用 201 次
