Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages
Tejas I. Dhamecha, V. Rudra Murthy, Samarth Bharadwaj, Karthik Sankaranarayanan, Pushpak Bhattacharyya
Abstract
We explore the impact of leveraging the relatedness of languages that belong to the same family in NLP models using multilingual fine-tuning. We hypothesize and validate that multilingual fine-tuning of pre-trained language models can yield better performance on downstream NLP applications, compared to models fine-tuned on individual languages. A first of its kind detailed study is presented to track performance change as languages are added to a base language in a graded and greedy (in the sense of best boost of performance) manner; which reveals that careful selection of subset of related languages can significantly improve performance than utilizing all related languages. The Indo-Aryan (IA) language family is chosen for the study, the exact languages being Bengali, Gujarati, Hindi, Marathi, Oriya, Punjabi and Urdu. The script barrier is crossed by simple rule-based transliteration of the text of all languages to Devanagari. Experiments are performed on mBERT, IndicBERT, MuRIL and two RoBERTa-based LMs, the last two being pre-trained by us. Low resource languages, such as Oriya and Punjabi, are found to be the largest beneficiaries of multilingual fine-tuning. Textual Entailment, Entity Classification, Section Title Prediction, tasks of IndicGLUE and POS tagging form our test bed. Compared to monolingual fine tuning we get relative performance improvement of up to 150% in the downstream tasks. The surprise take-away is that for any language there is a particular combination of other languages which yields the best performance, and any additional language is in fact detrimental.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ad50249-42db-43b0-883c-4a1b7ea51f70Cited by top-tier papers4
- Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra et al.ACL 2023 · 24 citations
- TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language ModelsYihong Liu, Chunlan Ma, Haotian Ye, Hinrich SchützeACL 2024 · 3 citations
- Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource LanguagesWenhao Zhuang, Yuan Sun, Xiaobing ZhaoACL 2025 · 1 citation
- Definition Generation for Word Meaning Modeling: Monolingual, Multilingual, and Cross-Lingual PerspectivesFrancesco Periti, Roksana Goworek, Haim Dubossarsky, Nina TahmasebiEMNLP 2025
Builds on6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified FrameworkZirui Wang, Jiateng Xie, Ruochen Xu, Yiming Yang et al.ICLR 2020 · 84 citations
Related papers
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 4 citations
- When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource LanguagesTyler A. Chang, Catherine Arnett, Zhuowen Tu, Ben BergenEMNLP 2024 · 12 citations
- ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicMuhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah NagoudiACL 2021
- Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal et al.ACL 2023 · 41 citations
- Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource LanguagesOlga Pelloni, Anastassia Shaitarova, Tanja SamardzicEMNLP 2022 · 4 citations
