Benchmarking Scalable Methods for Streaming Cross Document Entity Coreference
Robert L. Logan IV, Andrew McCallum, Sameer Singh, Daniel M. Bikel
Abstract
Streaming cross document entity coreference (CDC) systems disambiguate mentions of named entities in a scalable manner via incremental clustering. Unlike other approaches for named entity disambiguation (e.g., entity linking), streaming CDC allows for the disambiguation of entities that are unknown at inference time. Thus, it is well-suited for processing streams of data where new entities are frequently introduced. Despite these benefits, this task is currently difficult to study, as existing approaches are either evaluated on datasets that are no longer available, or omit other crucial details needed to ensure fair comparison. In this work, we address this issue by compiling a large benchmark adapted from existing free datasets, and performing a comprehensive evaluation of a number of novel and existing baseline models. 1 We investigate: how to best encode mentions, which clustering algorithms are most effective for grouping mentions, how models transfer to different domains, and how bounding the number of mentions tracked during inference impacts performance. Our results show that the relative performance of neural and feature-based mention encoders varies across different domains, and in most cases the best performance is achieved using a combination of both approaches. We also find that performance is minimally impacted by limiting the number of tracked mentions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7858e91-0f5d-4b4f-a4d9-3bced69a23daCited by top-tier papers3
- EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and IndexingNora Kassner, Fabio Petroni, Mikhail Plekhanov, Sebastian Riedel et al.EMNLP 2022 · 6 citations
- Sentence-Incremental Neural Coreference ResolutionMatt Grenander, Shay B. Cohen, Mark SteedmanEMNLP 2022 · 4 citations
- AcX: System, Techniques, and Experiments for Acronym ExpansionJoão L. M. Pereira, João Casanova, Helena Galhardas, Dennis E. ShashaVLDB 2022 · 2 citations
Builds on2
Related papers
- Contrastive Entity Coreference and Disambiguation for Historical TextsAbhishek Arora, Emily Silcock, Melissa Dell, Leander HeldringEMNLP 2024 · 1 citation
- Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document StreamsYukyung Lee, Yebin Lim, Woojun Jung, Wonjun Choi et al.KDD 2026
- xCoRe: Cross-context Coreference ResolutionGiuliano Martinelli, Bruno Gatti, Roberto NavigliEMNLP 2025
- Event Coreference Data (Almost) for Free: Mining Hyperlinks from Online NewsMichael Bugert, Iryna GurevychEMNLP 2021 · 4 citations
- Employing Discourse Coherence Enhancement to Improve Cross-Document Event and Entity Coreference ResolutionXinyu Chen, Peifeng Li, Qiaoming ZhuACL 2025
