Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages
Olga Pelloni, Anastassia Shaitarova, Tanja Samardzic
Abstract
Pre-trained multilingual models, such as mBERT, XLM-R and mT5, are used to improve the performance on various tasks in low-resource languages via cross-lingual transfer. In this framework, English is usually seen as the most natural choice for a transfer language (for fine-tuning or continued training of a multilingual pre-trained model), but it has been revealed recently that this is often not the best choice. The success of cross-lingual transfer seems to depend on some properties of languages, which are currently hard to explain. Successful transfer often happens between unrelated languages and it often cannot be explained by data-dependent factors.In this study, we show that languages written in non-Latin and non-alphabetic scripts (mostly Asian languages) are the best choices for improving performance on the task of Masked Language Modelling (MLM) in a diverse set of 30 low-resource languages and that the success of the transfer is well predicted by our novel measure of Subword Evenness (SuE). Transferring language models over the languages that score low on our measure results in the lowest average perplexity over target low-resource languages. Our correlation coefficients obtained with three different pre-trained multilingual models are consistently higher than all the other predictors, including text-based measures (type-token ratio, entropy) and linguistically motivated choice (genealogical and typological proximity).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c088444-a783-47b5-915a-d62d0f56f2deBuilds on7
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning TreatmentZirui Wang, Zachary C. Lipton, Yulia TsvetkovEMNLP 2020 · 72 citations
Related papers
- Cross-Lingual Ability of Multilingual Masked Language Models: A Study of Language StructureYuan Chai, Yaobo Liang, Nan DuanACL 2022
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related LanguagesVaidehi Patil, Partha P. Talukdar, Sunita SarawagiACL 2022 · 39 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
- Refining Low-Resource Unsupervised Translation by Language Disentanglement of Multilingual Translation ModelXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2022 · 6 citations
