MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition
David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula
摘要
African languages are spoken by over a billion people, but are underrepresented in NLP research and development. The challenges impeding progress include the limited availability of annotated datasets, as well as a lack of understanding of the settings where current methods are effective. In this paper, we make progress towards solutions for these challenges, focusing on the task of named entity recognition (NER). We create the largest human-annotated NER dataset for 20 African languages, and we study the behavior of state-of-the-art cross-lingual transfer methods in an Africa-centric setting, demonstrating that the choice of source language significantly affects performance. We show that choosing the best transfer language improves zero-shot F1 scores by an average of 14 points across 20 languages compared to using English. Our results highlight the need for benchmark datasets and models that cover typologically-diverse African languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Improving Language Plasticity via Pretraining with Active ForgettingYihong Chen, Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani 等NeurIPS 2023 · 被引用 49 次
- Constrained Decoding for Cross-lingual Label ProjectionDuong Minh Le, Yang Chen, Alan Ritter, Wei XuICLR 2024 · 被引用 14 次
- INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African LanguagesHao Yu, Jesujoba Oluwadara Alabi, Andiswa Bukula, Jian Yun Zhuang 等ACL 2025 · 被引用 8 次
- SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian LanguagesHoly Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V. Miranda 等EMNLP 2024 · 被引用 8 次
- CoLaDa: A Collaborative Label Denoising Framework for Cross-lingual Named Entity RecognitionTingting Ma, Qianhui Wu, Huiqiang Jiang, Börje Karlsson 等ACL 2023 · 被引用 5 次
它引用的顶会 Paper16
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
相关 Paper
- A Comparison of Architectures and Pretraining Methods for Contextualized Multilingual Word EmbeddingsNiels van der Heijden, Samira Abnar, Ekaterina ShutovaAAAI 2020 · 被引用 16 次
- Zero-Resource Cross-Lingual Named Entity RecognitionM. Saiful Bari, Shafiq R. Joty, Prathyusha JwalapuramAAAI 2020 · 被引用 55 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- Sources of Transfer in Multilingual Named Entity RecognitionDavid Mueller, Nicholas Andrews, Mark DredzeACL 2020 · 被引用 2 次
- MCL-NER: Cross-Lingual Named Entity Recognition via Multi-View Contrastive LearningYing Mo, Jian Yang, Jiahao Liu, Qifan Wang 等AAAI 2024 · 被引用 42 次
