SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian Languages
Prachuryya Kaushik, Ashish Anand
Abstract
We introduce SampurNER, a fine-grained named entity recognition (FgNER) dataset encompassing all 22 scheduled Indian languages spoken by more than two billion people across various countries. While manual annotation for FgNER resources is often labor-intensive and expensive, distant supervision methods have been employed as a viable solution. However, such datasets are often noisy, with entity mentions tagged with multiple types, requiring computationally intensive noise-aware models for effective FgNER. Moreover, resources for both coarse-grained and fine-grained named entity recognition tasks in Indian languages remain scarce. To address this, we propose an entity-anchored machine translation (EaMaTa) framework that leverages the largest manually annotated English FgNER dataset, FewNERD, to create a large-scale FgNER dataset in 22 languages. On average, the dataset comprises over 153k sentences, 354k entities, and 3.3M tokens in each language. The languages covered are: Assamese (as), Bengali (bn), Bodo (brx), Dogri (doi), Gujarati (gu), Hindi (hi), Kannada (kn), Kashmiri (ks), Konkani (gom), Maithili (mai), Malayalam (ml), Manipuri (mni), Marathi (mr), Nepali (ne), Odia (or), Punjabi (pa), Sanskrit (sa), Santali (sat), Sindhi (sd), Tamil (ta), Telugu (te), and Urdu (ur). Various rigorous analyses and human evaluations confirm the high quality of the dataset and demonstrate the effectiveness of the entity-anchored machine translation framework with up to 9% increase in F1-score against the current state-of-the-art. Additionally, we extend our analysis to zero-shot, multilingual, and cross-lingual settings, investigating the influence of language family and script similarity on cross-lingual FgNER performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal et al.ACL 2023 · 41 citations
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra et al.ACL 2023 · 24 citations
- XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationSebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant et al.EMNLP 2021 · 10 citations
- Adversity-aware Few-shot Named Entity Recognition via Augmentation LearningLi Huang, Haowen Liu, Qiang Gao, Jiajing Yu et al.AAAI 2025 · 1 citation
Related papers
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- Few-NERD: A Few-shot Named Entity Recognition DatasetNing Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang et al.ACL 2021
- Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource LanguagesXu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu et al.ACL 2022
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan et al.ACL 2023 · 8 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
