Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities
Sina Ahmadi, Antonios Anastasopoulos
摘要
The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language. This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script. Using synthetic data with various levels of noise and a transformerbased model, we demonstrate that the problem can be effectively remediated. We conduct a small-scale evaluation of real data as well. Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez ScriptHellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew, Hizkiel Mitiku Alemayehu 等EMNLP 2025 · 被引用 1 次
- Phonetic and Visual Priors for Decipherment of Informal RomanizationMaria Ryskina, Matthew R. Gormley, Taylor Berg-KirkpatrickACL 2020
- Arabic Diacritics in the Wild: Exploiting Opportunities for Improved DiacritizationSalman Elgamal, Ossama Obeid, Mhd Tameem Kabbani, Go Inoue 等ACL 2024 · 被引用 1 次
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 被引用 3 次
- RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via RomanizationJaavid Aktar Husain, Raj Dabre, Aswanth M., Jay Gala 等ACL 2024 · 被引用 1 次
