Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
Refael Shaked Greenfeld, Reut Tsarfaty
Abstract
Coreference Resolution (CR) is a fundamental NLP task critical for long-form tasks as information extraction, summarization, and many business applications. However, CR methods originally designed for English struggle with Morphologically Rich Languages (MRLs), where mention boundaries do not necessarily align with word boundaries, and a single token may consist of multiple anaphors. CR modeling and evaluation protocols standardly assume that, as in English, words and mentions mostly align. However, this assumption breaks down in MRLs, particularly in the context of LLMs'raw-text processing and end-to-end tasks. To assess and address this challenge, we introduce KibutzR, the first comprehensive CR dataset for Modern Hebrew, an MRL rich with complex words and pronominal clitics. We deliver an annotated dataset that identifies mentions at word, sub-word and multi-word levels, and propose an evaluation protocol that directly addresses word/morpheme boundary discrepancies. Our experiments show that contemporary LLMs perform significantly worse on Hebrew than on English, and that performance degrades on raw unsegmented text. Crucially, we show an inverse performance-trend in Hebrew relative to English, where smaller encoders perform far better than contemporary decoder models, leaving ample space for investigation and improvement. We deliver a new benchmark for Hebrew coreference resolution and a segmentation-aware evaluation protocol to inform future work on other MRLs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter RefinementKangyang Luo, Yuzhuo Bai, Shuzheng Si, Cheng Gao et al.ACL 2026 · 1 citation
- Probing for Referential Information in Language ModelsIonut-Teodor Sorodoc, Kristina Gulordava, Gemma BoledaACL 2020 · 31 citations
- From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?Reut Tsarfaty, Dan Bareket, Stav Klein, Amit SekerACL 2020 · 2 citations
- xCoRe: Cross-context Coreference ResolutionGiuliano Martinelli, Bruno Gatti, Roberto NavigliEMNLP 2025
- Testing Coreference Resolution Systems without Labeled Test SetsJialun Cao, Yaojie Lu, Ming Wen, Shing-Chi CheungFSE 2023 · 3 citations
