Don't Stop Fine-Tuning: On Training Regimes for Few-Shot Cross-Lingual Transfer with Multilingual Language Models
Fabian David Schmidt, Ivan Vulic, Goran Glavas
Abstract
A large body of recent work highlights the fallacies of zero-shot cross-lingual transfer (ZS-XLT) with large multilingual language models. Namely, their performance varies substantially for different target languages and is the weakest where needed the most: for low-resource languages distant to the source language. One remedy is few-shot transfer (FS-XLT), where leveraging only a few task-annotated instances in the target language(s) may yield sizable performance gains. However, FS-XLT also succumbs to large variation, as models easily overfit to the small datasets. In this work, we present a systematic study focused on a spectrum of FS-XLT fine-tuning regimes, analyzing key properties such as effectiveness, (in)stability, and modularity. We conduct extensive experiments on both higher-level (NLI, paraphrasing) and lower-level tasks (NER, POS), presenting new FS-XLT strategies that yield both improved and more stable FS-XLT across the board. Our findings challenge established FS-XLT methods: e.g., we propose to replace sequential fine-tuning with joint fine-tuning on source and target language instances, offering consistent gains with different number of shots (including resource-rich scenarios). We also show that further gains can be achieved with multi-stage FS-XLT training in which joint multilingual fine-tuning precedes the bilingual source-target specialization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio et al.EMNLP 2024 · 10 citations
- A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue SystemsSongbo Hu, Han Zhou, Moy Yuan, Milan Gritta et al.EMNLP 2023 · 6 citations
- Teaching LLMs to Abstain across Languages via Multilingual FeedbackShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding et al.EMNLP 2024 · 4 citations
- Free Lunch: Robust Cross-Lingual Transfer via Model Checkpoint AveragingFabian David Schmidt, Ivan Vulic, Goran GlavasACL 2023 · 3 citations
- Efficiently Maintaining the Multilingual Capacity of MCLIP in Downstream Cross-Modal Retrieval TasksFengmao Lyu, Jitong Lei, Guosheng Lin, Desheng Zheng et al.NeurIPS 2025 · 2 citations
Builds on15
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
Related papers
- A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots MattersMengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulic et al.ACL 2021
- Less-forgetting Multi-lingual Fine-tuningYuren Mao, Yaobo Liang, Nan Duan, Haobo Wang et al.NeurIPS 2022 · 10 citations
- ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersVipul Rathore, Rajdeep Dhingra, Parag Singla, MausamEMNLP 2023
- Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsKabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit ChoudhuryACL 2022
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
