Improving Informally Romanized Language Identification
Adrian Benton, Alexander Gutkin, Christo Kirov, Brian Roark
Abstract
The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such romanization renders languages that are normally easily distinguished due to being written in different scripts - Hindi and Urdu, for example - highly confusable. In this work, we increase language identification (LID) accuracy for romanized text by improving the methods used to synthesize training sets. We find that training on synthetic samples which incorporate natural spelling variation yields higher LID system accuracy than including available naturally occurring examples in the training set, or even training higher capacity models. We demonstrate new state-of-the-art LID performance on romanized text from 20 Indic languages in the Bhasha-Abhijnaanam evaluation set (Madhani et al., 2023a), improving test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% using a linear classifier trained solely on synthetic data and 88.2% when also training on available harvested text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- AfroLID: A Neural Language Identification Tool for African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba InciarteEMNLP 2022 · 13 citations
- One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in IndonesiaAlham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya et al.ACL 2022
Related papers
- RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via RomanizationJaavid Aktar Husain, Raj Dabre, Aswanth M., Jay Gala et al.ACL 2024 · 1 citation
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 3 citations
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic LanguagesGuduru Manoj, Neel Prabhanjan Rachamalla, Ashish Kulkarni, Gautam Rajeev et al.AAAI 2026 · 2 citations
- Phonetic and Visual Priors for Decipherment of Informal RomanizationMaria Ryskina, Matthew R. Gormley, Taylor Berg-KirkpatrickACL 2020
- NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous ScriptsMuhammad Farid Adilazuarda, Musa Izzanardi Wijanarko, Lucky Susanto, Khumaisa Nur'aini et al.ACL 2025
