Naamapadam: A Large-Scale Named Entity Annotated Data for Indic Languages
Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, Pratyush Kumar, V. Rudra Murthy, Anoop Kunchukuttan
Abstract
We present, Naamapadam, the largest publicly available Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families. The dataset contains more than 400k sentences annotated with a total of at least 100k entities from three standard entity categories (Person, Location, and, Organization) for 9 out of the 11 languages. The training dataset has been automatically created from the Samanantar parallel corpus by projecting automatically tagged entities from an English sentence to the corresponding Indian language translation. We also create manually annotated testsets for 9 languages. We demonstrate the utility of the obtained dataset on the Naamapadam-test dataset. We also release In-dicNER, a multilingual IndicBERT model finetuned on Naamapadam training set. IndicNER achieves an F1 score of more than 80 for 7 out of 9 test languages. The dataset and models are available under open-source licences at https: //ai4bharat.iitm.ac.in/naamapadam .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal et al.ACL 2023 · 41 citations
- Multilingual Pretraining for Pixel Language ModelsIlker Kesen, Jonas F. Lotz, Ingo Ziegler, Phillip Rust et al.EMNLP 2025 · 1 citation
- Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian LanguagesAmir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri KotoACL 2026
- SampurNER: Fine-Grained Named Entity Recognition Dataset for 22 Indian LanguagesPrachuryya Kaushik, Ashish AnandAAAI 2026
Builds on11
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Dice Loss for Data-imbalanced NLP TasksXiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang et al.ACL 2020 · 575 citations
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
Related papers
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li et al.EMNLP 2025
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary et al.ACL 2025 · 5 citations
- Code and Named Entity Recognition in StackOverflowJeniya Tabassum, Mounica Maddela, Wei Xu, Alan RitterACL 2020 · 9 citations
