Comparative evaluation of boundary-relaxed annotation for Entity Linking performance
Gabriel Herman Bernardim Andrade, Shuntaro Yada, Eiji Aramaki
Abstract
Entity Linking performance has a strong reliance on having a large quantity of high-quality annotated training data available. Yet, manual annotation of named entities, especially their boundaries, is ambiguous, error-prone, and raises many inconsistencies between annotators. While imprecise boundary annotation can degrade a model's performance, there are applications where accurate extraction of entities' surface form is not necessary. For those cases, a lenient annotation guideline could relieve the annotators' workload and speed up the process. This paper presents a case study designed to verify the feasibility of such annotation process and evaluate the impact of boundary-relaxed annotation in an Entity Linking pipeline. We first generate a set of noisy versions of the widely used AIDA CoNLL-YAGO dataset by expanding the boundaries subsets of annotated entity mentions and then train three Entity Linking models on this data and evaluate the relative impact of imprecise annotation on entity recognition and disambiguation performances. We demonstrate that the magnitude of effects caused by noise in the Named Entity Recognition phase is dependent on both model complexity and noise ratio, while Entity Disambiguation components are susceptible to entity boundary imprecision due to strong vocabulary dependency. Original annotation Noisy annotation CRICKET -LEICESTERSHIRE TAKE OVER AT TOP AFTER INNINGS VICTORY. LONDON 1996-08-30 West Indian all-rounder Phil Simmons took four for 38 on Friday as Leicestershire beat Somerset by an innings and 39 runs in two days to take over at the head of the county championship. CRICKET -LEICESTERSHIRE TAKE OVER AT TOP AFTER INNINGS VICTORY. LONDON 1996-08-30 West Indian all-rounder Phil Simmons took four for 38 on Friday as Leicestershire beat Somerset by an innings and 39 runs in two days to take over at the head of the county championship.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- Boundary Smoothing for Named Entity RecognitionEnwei Zhu, Jinpeng LiACL 2022 · 92 citations
- EntQA: Entity Linking as Question AnsweringWenzheng Zhang, Wenyue Hua, Karl StratosICLR 2022 · 64 citations
- Weakly Supervised Named Entity Tagging with Learnable Logical RulesJiacheng Li, Haibo Ding, Jingbo Shang, Julian J. McAuley et al.ACL 2021
Related papers
- CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetSusanna Rücker, Alan AkbikEMNLP 2023 · 3 citations
- From Zero to Hero: Human-In-The-Loop Entity Linking in Low Resource DomainsJan-Christoph Klie, Richard Eckart de Castilho, Iryna GurevychACL 2020 · 42 citations
- Fine-Grained Entity Typing for Domain Independent Entity LinkingYasumasa Onoe, Greg DurrettAAAI 2020 · 94 citations
- Robustness Evaluation of Entity Disambiguation Using Prior Probes: the Case of Entity OvershadowingVera Provatorova, Samarth Bhargav, Svitlana Vakulenko, Evangelos KanoulasEMNLP 2021 · 7 citations
- Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsJian Liu, Weichang Liu, Yufeng Chen, Jinan Xu et al.EMNLP 2023 · 3 citations
