Exploiting Language Relatedness for Low Web-Resource Language Model Adaptation: An Indic Languages Study
Yash Khemchandani, Sarvesh Mehtani, Vaidehi Patil, Abhijeet Awasthi, Partha P. Talukdar, Sunita Sarawagi
Abstract
Recent research in multilingual language models (LM) has demonstrated their ability to effectively handle multiple languages in a single model. This holds promise for low web-resource languages (LRL) as multilingual models can enable transfer of supervision from high resource languages to LRLs. However, incorporating a new language in an LM still remains a challenge, particularly for languages with limited corpora and in unseen scripts. In this paper we argue that relatedness among languages in a language family may be exploited to overcome some of the corpora limitations of LRLs, and propose RelateLM. We focus on Indian languages, and exploit relatedness along two dimensions: (1) script (since many Indic scripts originated from the Brahmic script), and (2) sentence structure. RelateLM uses transliteration to convert the unseen script of limited LRL text into the script of a Related Prominent Language (RPL) (Hindi in our case). While exploiting similar sentence structures, RelateLM utilizes readily available bilingual dictionaries to pseudo translate RPL text into LRL corpora. Experiments on multiple real-world benchmark datasets provide validation to our hypothesis that using a related language as pivot, along with transliteration and pseudo translation based data augmentation, can be an effective way to adapt LMs for LRLs, rather than direct training or pivoting through English.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Expanding Pretrained Models to Thousands More Languages via Lexicon-based AdaptationXinyi Wang, Sebastian Ruder, Graham NeubigACL 2022 · 73 citations
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal et al.ACL 2023 · 41 citations
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesVaidehi Patil, Partha P. Talukdar, Sunita SarawagiACL 2022 · 39 citations
- RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via RomanizationJaavid Aktar Husain, Raj Dabre, Aswanth M., Jay Gala et al.ACL 2024 · 1 citation
- Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationJaeseong Lee, Dohyeon Lee, Seung-won HwangAAAI 2023 · 1 citation
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Multilingual Alignment of Contextual Word RepresentationsSteven Cao, Nikita Kitaev, Dan KleinICLR 2020 · 211 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource LanguagesWenhao Zhuang, Yuan Sun, Xiaobing ZhaoACL 2025 · 1 citation
- Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan LanguagesTejas I. Dhamecha, V. Rudra Murthy, Samarth Bharadwaj, Karthik Sankaranarayanan et al.EMNLP 2021 · 18 citations
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 4 citations
- Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource LanguagesOlga Pelloni, Anastassia Shaitarova, Tanja SamardzicEMNLP 2022 · 4 citations
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 3 citations
