Benchmarking Scalable Methods for Streaming Cross Document Entity Coreference
Robert L. Logan IV, Andrew McCallum, Sameer Singh, Daniel M. Bikel
摘要
Streaming cross document entity coreference (CDC) systems disambiguate mentions of named entities in a scalable manner via incremental clustering. Unlike other approaches for named entity disambiguation (e.g., entity linking), streaming CDC allows for the disambiguation of entities that are unknown at inference time. Thus, it is well-suited for processing streams of data where new entities are frequently introduced. Despite these benefits, this task is currently difficult to study, as existing approaches are either evaluated on datasets that are no longer available, or omit other crucial details needed to ensure fair comparison. In this work, we address this issue by compiling a large benchmark adapted from existing free datasets, and performing a comprehensive evaluation of a number of novel and existing baseline models. 1 We investigate: how to best encode mentions, which clustering algorithms are most effective for grouping mentions, how models transfer to different domains, and how bounding the number of mentions tracked during inference impacts performance. Our results show that the relative performance of neural and feature-based mention encoders varies across different domains, and in most cases the best performance is achieved using a combination of both approaches. We also find that performance is minimally impacted by limiting the number of tracked mentions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and IndexingNora Kassner, Fabio Petroni, Mikhail Plekhanov, Sebastian Riedel 等EMNLP 2022 · 被引用 6 次
- Sentence-Incremental Neural Coreference ResolutionMatt Grenander, Shay B. Cohen, Mark SteedmanEMNLP 2022 · 被引用 4 次
- AcX: System, Techniques, and Experiments for Acronym ExpansionJoão L. M. Pereira, João Casanova, Helena Galhardas, Dennis E. ShashaVLDB 2022 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Contrastive Entity Coreference and Disambiguation for Historical TextsAbhishek Arora, Emily Silcock, Melissa Dell, Leander HeldringEMNLP 2024 · 被引用 1 次
- Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document StreamsYukyung Lee, Yebin Lim, Woojun Jung, Wonjun Choi 等KDD 2026
- xCoRe: Cross-context Coreference ResolutionGiuliano Martinelli, Bruno Gatti, Roberto NavigliEMNLP 2025
- Event Coreference Data (Almost) for Free: Mining Hyperlinks from Online NewsMichael Bugert, Iryna GurevychEMNLP 2021 · 被引用 4 次
- Employing Discourse Coherence Enhancement to Improve Cross-Document Event and Entity Coreference ResolutionXinyu Chen, Peifeng Li, Qiaoming ZhuACL 2025
