AfroLID: A Neural Language Identification Tool for African Languages
Ife Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba Inciarte
Abstract
Language identification (LID) is a crucial precursor for NLP, especially for mining web data. Problematically, most of the world's 7000+ languages today are not covered by LID technologies. We address this pressing issue for Africa by introducing AfroLID, a neural LID toolkit for 517 African languages and varieties. AfroLID exploits a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems. When evaluated on our blind Test set, AfroLID achieves 95.89 F 1 -score. We also compare AfroLID to five existing LID tools that each cover a small number of African languages, finding it to outperform them on most languages. We further show the utility of AfroLID in the wild by testing it on the acutely under-served Twitter domain. Finally, we offer a number of controlled case studies and perform a linguistically-motivated error analysis that allow us to both showcase AfroLID's powerful capabilities and limitations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 431ad3cd-610d-4dd5-a671-33a9efd4294bCited by top-tier papers4
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera et al.ACL 2026 · 5 citations
- LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ LanguagesMilind Agarwal, Md Mahfuz Ibn Alam, Antonios AnastasopoulosEMNLP 2023 · 2 citations
- Improving Informally Romanized Language IdentificationAdrian Benton, Alexander Gutkin, Christo Kirov, Brian RoarkEMNLP 2025 · 1 citation
- Voice of a Continent: Mapping Africa's Speech Technology FrontierAbdelRahim A. Elmadany, Sang Yun Kwon, Hawau Olamide Toyin, Alcides Alcoba Inciarte et al.EMNLP 2025
Builds on7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- AraT5: Text-to-Text Transformers for Arabic Language GenerationEl Moatez Billah Nagoudi, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2022 · 175 citations
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 72 citations
- CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsAhmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp KoehnEMNLP 2020 · 6 citations
- ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicMuhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah NagoudiACL 2021
Related papers
- AfriSenti: A Twitter Sentiment Analysis Benchmark for African LanguagesShamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum et al.EMNLP 2023 · 33 citations
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
- The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLPSheriff Issaka, Keyi Wang, Yinka Ajibola, Oluwatumininu Samuel-Ipaye et al.ACL 2026
- MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity RecognitionDavid Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani et al.EMNLP 2022 · 46 citations
- Identifying Open Challenges in Language IdentificationRob van der GootACL 2025
