SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages
Prachuryya Kaushik, Ashish Anand
摘要
We introduce SampurNER, a fine-grained named entity recognition (FgNER) dataset encompassing all 22 scheduled Indian languages spoken by more than two billion people across various countries. While manual annotation for FgNER resources is often labor-intensive and expensive, distant supervision methods have been employed as a viable solution. However, such datasets are often noisy, with entity mentions tagged with multiple types, requiring computationally intensive noise-aware models for effective FgNER. Moreover, resources for both coarse-grained and fine-grained named entity recognition tasks in Indian languages remain scarce. To address this, we propose an entity-anchored machine translation (EaMaTa) framework that leverages the largest manually annotated English FgNER dataset, FewNERD, to create a large-scale FgNER dataset in 22 languages. On average, the dataset comprises over 153k sentences, 354k entities, and 3.3M tokens in each language. The languages covered are: Assamese (as), Bengali (bn), Bodo (brx), Dogri (doi), Gujarati (gu), Hindi (hi), Kannada (kn), Kashmiri (ks), Konkani (gom), Maithili (mai), Malayalam (ml), Manipuri (mni), Marathi (mr), Nepali (ne), Odia (or), Punjabi (pa), Sanskrit (sa), Santali (sat), Sindhi (sd), Tamil (ta), Telugu (te), and Urdu (ur). Various rigorous analyses and human evaluations confirm the high quality of the dataset and demonstrate the effectiveness of the entity-anchored machine translation framework with up to 9% increase in F1-score against the current state-of-the-art. Additionally, we extend our analysis to zero-shot, multilingual, and cross-lingual settings, investigating the influence of language family and script similarity on cross-lingual FgNER performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal 等ACL 2023 · 被引用 41 次
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra 等ACL 2023 · 被引用 24 次
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant 等EMNLP 2021 · 被引用 10 次
- Adversity-aware Few-shot Named Entity Recognition via Augmentation LearningLi Huang, Haowen Liu, Qiang Gao, Jiajing Yu 等AAAI 2025 · 被引用 1 次
相关 Paper
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani 等EMNLP 2022 · 被引用 46 次
- Few-NERD: A Few-shot Named Entity Recognition DatasetNing Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang 等ACL 2021
- Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource LanguagesXu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu 等ACL 2022
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan 等ACL 2023 · 被引用 8 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
