How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
Rochelle Choenni, Dan Garrette, Ekaterina Shutova
Abstract
Multilingual language models (MLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languages’ data. Impressive performance in zero-shot cross-lingual transfer shows that these models are able to exploit this property. Yet, it remains unclear to what extent, and under which conditions, languages rely on each other’s data. To answer this question, we use TracIn (Pruthi et al., 2020), a training data attribution (TDA) method, to retrieve training samples from multilingual data that are most influential for test predictions in a given language. This allows us to analyse cross-lingual sharing mechanisms of MLMs from a new perspective. While previous work studied cross-lingual sharing at the model parameter level, we present the first approach to study it at the data level. We find that MLMs rely on data from multiple languages during fine-tuning and this reliance increases as fine-tuning progresses. We further find that training samples from other languages can both reinforce and complement the knowledge acquired from data of the test language itself.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b514e354-357a-449d-942c-589007a651eeCited by top-tier papers5
- Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?Dawei Zhu, Pinzhen Chen, Miaoran Zhang, Barry Haddow et al.EMNLP 2024 · 3 citations
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models PretrainingPing Guo, Yubing Ren, Binbin Liu, Fengze Liu et al.NeurIPS 2025 · 2 citations
- Shared Path: Unraveling Memorization in Multilingual LLMs through Language SimilaritiesXiaoyu Luo, Yiyi Chen, Johannes Bjerva, Qiongxiu LiEMNLP 2025 · 1 citation
- The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language ModelJiawei Chen, Wentao Chen, Jing Su, Jingjing Xu et al.ICLR 2025
- Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language ModelsLucas Bandarkar, Benjamin Muller, Pritish Yuvraj, Rui Hou et al.ICLR 2025
Builds on13
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 91 citations
Related papers
- Analysis of Multi-Source Language Training in Cross-Lingual TransferSeong Hoon Lim, Taejun Yun, Jinhyeon Kim, Jihun Choi et al.ACL 2024 · 1 citation
- The Echoes of Multilinguality: Tracing Cultural Value Shifts during Language Model Fine-tuningRochelle Choenni, Anne Lauscher, Ekaterina ShutovaACL 2024 · 2 citations
- Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationJunpeng Liu, Kaiyu Huang, Jiuyi Li, Huan Liu et al.EMNLP 2022 · 5 citations
- Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsKabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit ChoudhuryACL 2022
- Effective Fine-Tuning Methods for Cross-lingual AdaptationTao Yu, Shafiq R. JotyEMNLP 2021 · 7 citations
