Identifying Open Challenges in Language Identification
Rob van der Goot
Abstract
Automatic language identification is a core problem of many Natural Language Processing (NLP) pipelines. A wide variety of architectures and benchmarks have been proposed with often near-perfect performance. Although previous studies have focused on certain challenging setups (i.e. cross-domain, short inputs), a systematic comparison is missing. We propose a benchmark that allows us to test for the effect of input size, training data size, domain, number of languages, scripts, and language families on performance. We evaluate five popular models on this benchmark and identify which open challenges remain for this task as well as which architectures achieve robust performance. We find that cross-domain setups are the most challenging (although arguably most relevant), and that number of languages, variety in scripts, and variety in language families have only a small impact on performance. We also contribute practical takeaways: training with 1,000 instances per language and a maximum input length of 100 characters is enough for robust language identification. Based on our findings, we train an accurate (94.41%) multidomain language identification model on 2,034 languages, for which we also provide an analysis of the remaining errors. 1References
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7fb76d3-b132-46c9-bf49-b24b64d0d83bCited by top-tier papers2
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera et al.ACL 2026 · 5 citations
- What Language is This? Ask Your Tokenizer.Clara Meister, Ahmetcan Yavuz, Pietro Lesci, Tiago PimentelICML 2026 · 1 citation
Related papers
- LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ LanguagesMilind Agarwal, Md Mahfuz Ibn Alam, Antonios AnastasopoulosEMNLP 2023 · 2 citations
- DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related LanguagesFahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja et al.ACL 2024 · 10 citations
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng et al.ACL 2025 · 3 citations
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 3 citations
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng et al.ICLR 2026 · 27 citations
