CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja
Abstract
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID's value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd2f65ee-8d08-43a6-bf7b-57d88fbf91fcBuilds on4
- AfroLID: A Neural Language Identification Tool for African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba InciarteEMNLP 2022 · 13 citations
- ALDi: Quantifying the Arabic Level of Dialectness of TextAmr Keleg, Sharon Goldwater, Walid MagdyEMNLP 2023 · 4 citations
- Identifying Open Challenges in Language IdentificationRob van der GootACL 2025
- CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WebHolger Schwenk, Guillaume Wenzek, Sergey Edunov, Edouard Grave et al.ACL 2021
Related papers
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza et al.ICLR 2026 · 22 citations
- What Language is This? Ask Your Tokenizer.Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago PimentelICML 2026 · 1 citation
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas et al.EMNLP 2023 · 25 citations
- Generative Language Models for Paragraph-Level Question GenerationAsahi Ushio, Fernando Alva-Manchego, José Camacho-ColladosEMNLP 2022 · 30 citations
