LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages
Milind Agarwal, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
Abstract
Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems. Language identification is a well-studied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, current systems cannot accurately identify most of the world’s 7000 languages. To tackle this bottleneck, we first compile a corpus, MCS-350, of 50K multilingual and parallel children’s stories in 350+ languages. MCS-350 can serve as a benchmark for language identification of short texts and for 1400+ new translation directions in low-resource Indian and African languages. Second, we propose a novel misprediction-resolution hierarchical model, LIMIT, for language identification that reduces error by 55% (from 0.71 to 0.32) on our compiled children’s stories dataset and by 40% (from 0.23 to 0.14) on the FLORES-200 benchmark. Our method can expand language identification coverage into low-resource languages by relying solely on systemic misprediction patterns, bypassing the need to retrain large models from scratch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 97 citations
- Balancing Training for Multilingual Neural Machine TranslationXinyi Wang, Yulia Tsvetkov, Graham NeubigACL 2020 · 74 citations
- AfroLID: A Neural Language Identification Tool for African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba InciarteEMNLP 2022 · 13 citations
- Contrastive Learning for Many-to-many Multilingual Neural Machine TranslationXiao Pan, Mingxuan Wang, Liwei Wu, Lei LiACL 2021
Related papers
- Identifying Open Challenges in Language IdentificationRob van der GootACL 2025
- ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language ModelsEmily Chang, Niyati BafnaACL 2026 · 1 citation
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- What Language is This? Ask Your Tokenizer.Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago PimentelICML 2026 · 1 citation
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera et al.ACL 2026 · 5 citations
