OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä, Constantine Lignos
Abstract
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, humanannotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER. We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER. We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task. OpenNER is released at https://github.com/bltlab/open-ner . 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ae8ad4c-eceb-4f6b-8de5-7fe67b529cbfBuilds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- Empirical Study of Zero-Shot NER with ChatGPTTingyu Xie, Qi Li, Jian Zhang, Yan Zhang et al.EMNLP 2023 · 50 citations
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- Building a User-Generated Content North-African Arabizi Treebank: Tackling HellDjamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral et al.ACL 2020 · 38 citations
Related papers
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li et al.EMNLP 2025
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra et al.ACL 2023 · 24 citations
- Sources of Transfer in Multilingual Named Entity RecognitionDavid Mueller, Nicholas Andrews, Mark DredzeACL 2020 · 2 citations
- NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial RegistriesSimona Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman et al.EMNLP 2024
- NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated DataSergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé et al.EMNLP 2024 · 29 citations
