CleanCoNLL: A Nearly Noise-Free Named Entity Recognition Dataset
Susanna Rücker, Alan Akbik
摘要
The CoNLL-03 corpus is arguably the most well-known and utilized benchmark dataset for named entity recognition (NER). However, prior works found significant numbers of annotation errors, incompleteness, and inconsistencies in the data. This poses challenges to objectively comparing NER approaches and analyzing their errors, as current state-of-the-art models achieve F1-scores that are comparable to or even exceed the estimated noise level in CoNLL-03. To address this issue, we present a comprehensive relabeling effort assisted by automatic consistency checking that corrects 7.0% of all labels in the English CoNLL-03. Our effort adds a layer of entity linking annotation both for better explainability of NER labels and as additional safeguard of annotation quality. Our experimental evaluation finds not only that state-of-the-art approaches reach significantly higher F1-scores (97.1%) on our data, but crucially that the share of correct predictions falsely counted as errors due to annotation noise drops from 47% to 6%. This indicates that our resource is well suited to analyze the remaining errors made by state-of-the-art models, and that the theoretical upper bound even on high resource, coarse-grained NER is not yet reached. To facilitate such analysis, we make CLEANCONLL publicly available to the research community 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionElena Merdjanovska, Ansar Aynetdinov, Alan AkbikEMNLP 2024 · 被引用 5 次
- VariErr NLI: Separating Annotation Error from Human Label VariationLeon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara PlankACL 2024
它引用的顶会 Paper8
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Rethinking Generalization of Neural Models: A Named Entity Recognition Case StudyJinlan Fu, Pengfei Liu, Qi ZhangAAAI 2020 · 被引用 79 次
- Learning from Noisy Labels for Entity-Centric Information ExtractionWenxuan Zhou, Muhao ChenEMNLP 2021 · 被引用 32 次
相关 Paper
- Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsJian Liu, Weichang Liu, Yufeng Chen, Jinan Xu 等EMNLP 2023 · 被引用 3 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- Learning "O" Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NERRuotian Ma, Xuanting Chen, Zhang Lin, Xin Zhou 等ACL 2023 · 被引用 12 次
- Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023?Shuheng Liu, Alan RitterACL 2023 · 被引用 9 次
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor 等EMNLP 2025 · 被引用 9 次
