Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
Sujoy Sarkar, Gourav Sarkar, Manoj Balaji Jagadeeshan, Jivnesh Sandhan, Amrith Krishna, Pawan Goyal
摘要
High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging. We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language. Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking. The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems. Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set. These results highlight the limitations of current approaches in resolving entities within such complex discourse. Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Entities as Experts: Sparse Memory Access with Entity SupervisionThibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi 等EMNLP 2020 · 被引用 39 次
- Dual Cache for Long Document Neural Coreference ResolutionQipeng Guo, Xiangkun Hu, Yue Zhang, Xipeng Qiu 等ACL 2023 · 被引用 4 次
- Contrastive Entity Coreference and Disambiguation for Historical TextsAbhishek Arora, Emily Silcock, Melissa Dell, Leander HeldringEMNLP 2024 · 被引用 1 次
- Evaluating Entity Disambiguation and the Role of Popularity in Retrieval-Based NLPAnthony Chen, Pallavi Gudipati, Shayne Longpre, Xiao Ling 等ACL 2021
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
相关 Paper
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra 等ACL 2023 · 被引用 24 次
- BOOKCOREF: Coreference Resolution at Book ScaleGiuliano Martinelli, Tommaso Bonomo, Pere-Lluís Huguet Cabot, Roberto NavigliACL 2025
- Multimodal Entity Linking: A New Dataset and A BaselineJingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang 等ACM MM 2021 · 被引用 41 次
- Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex TextRefael Shaked Greenfeld, Reut TsarfatyACL 2026
- GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed TextMichael Ginn, Lindia Tjuatja, Taiqi He, Enora Rice 等EMNLP 2024 · 被引用 1 次
