NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
Elena Merdjanovska, Ansar Aynetdinov, Alan Akbik
摘要
Available training data for named entity recognition (NER) often contains a significant percentage of incorrect labels for entity types and entity boundaries. Such label noise poses challenges for supervised learning and may significantly deteriorate model quality. To address this, prior work proposed various noise-robust learning approaches capable of learning from data with partially incorrect labels. These approaches are typically evaluated using simulated noise where the labels in a clean dataset are automatically corrupted. However, as we show in this paper, this leads to unrealistic noise that is far easier to handle than real noise caused by human error or semi-automatic annotation. To enable the study of the impact of various types of real noise, we introduce NOISEBENCH, an NER benchmark consisting of clean training data corrupted with 6 types of real noise, including expert errors, crowdsourcing errors, automatic annotation errors and LLM errors. We present an analysis that shows that real noise is significantly more challenging than simulated noise, and show that current state-of-the-art models for noise-robust learning fall far short of their achievable upper bound. We release NOISEBENCH for both English and German to the research community 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 被引用 241 次
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Memorisation versus Generalisation in Pre-trained Language ModelsMichael Tänzer, Sebastian Ruder, Marek ReiACL 2022 · 被引用 59 次
- Learning from Noisy Labels for Entity-Centric Information ExtractionWenxuan Zhou, Muhao ChenEMNLP 2021 · 被引用 32 次
- Analysing the Noise Model Error for Realistic Noisy Label DataMichael A. Hedderich, Dawei Zhu, Dietrich KlakowAAAI 2021 · 被引用 26 次
相关 Paper
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang 等EMNLP 2021 · 被引用 50 次
- CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetSusanna Rücker, Alan AkbikEMNLP 2023 · 被引用 3 次
- AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionRuipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu 等ICML 2026 · 被引用 12 次
- NAT: Noise-Aware Training for Robust Neural Sequence LabelingMarcin Namysl, Sven Behnke, Joachim KöhlerACL 2020
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 被引用 8 次
