Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities
Sina Ahmadi, Antonios Anastasopoulos
Abstract
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language. This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script. Using synthetic data with various levels of noise and a transformerbased model, we demonstrate that the problem can be effectively remediated. We conduct a small-scale evaluation of real data as well. Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6a58df0-a4a2-467f-94ba-4de72a597075Related papers
- A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez ScriptHellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew, Hizkiel Mitiku Alemayehu et al.EMNLP 2025 · 1 citation
- Phonetic and Visual Priors for Decipherment of Informal RomanizationMaria Ryskina, Matthew R. Gormley, Taylor Berg-KirkpatrickACL 2020
- Arabic Diacritics in the Wild: Exploiting Opportunities for Improved DiacritizationSalman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue et al.ACL 2024 · 1 citation
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 3 citations
- RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via RomanizationJaavid Aktar Husain, Raj Dabre, Aswanth M., Jay Gala et al.ACL 2024 · 1 citation
