OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä, Constantine Lignos
摘要
We present OpenNER 1.0, a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, humanannotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and multi-ontology NER. We provide baseline results using three pretrained multilingual language models and two large language models to compare the performance of recent models and facilitate future research in NER. We find that no single model is best in all languages and that significant work remains to obtain high performance from LLMs on the NER task. OpenNER is released at https://github.com/bltlab/open-ner . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen 等ICLR 2024 · 被引用 118 次
- Empirical Study of Zero-Shot NER with ChatGPTTingyu Xie, Qi Li, Jian Zhang, Yan Zhang 等EMNLP 2023 · 被引用 50 次
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani 等EMNLP 2022 · 被引用 46 次
- Building a User-Generated Content North-African Arabizi Treebank: Tackling HellDjamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral 等ACL 2020 · 被引用 38 次
相关 Paper
- DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity RecognitionHanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li 等EMNLP 2025
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra 等ACL 2023 · 被引用 24 次
- Sources of Transfer in Multilingual Named Entity RecognitionDavid Mueller, Nicholas Andrews, Mark DredzeACL 2020 · 被引用 2 次
- NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial RegistriesSimona Doneva, Tilia Ellendorff, Beate Sick, Jean-Philippe Goldman 等EMNLP 2024
- NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated DataSergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé 等EMNLP 2024 · 被引用 29 次
