The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani, Pontus Stenetorp, Jianfei Yang, Yao Lu
Abstract
Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely believed to enable these abilities, details of its contributions remain unclear. We investigate this question by pretraining models from scratch under controlled conditions, comparing the standard web corpus with a monolingual-only version that removes all multilingual documents. Despite constituting only 2% of the corpus, removing bilingual data causes translation performance to drop 56% in BLEU, while behaviour on cross-lingual QA and general reasoning tasks remains stable, with training curves largely overlapping the baseline. To understand this asymmetry, we categorize bilingual data into parallel (14%), code-switching (72%), and miscellaneous documents (14%) based on the semantic relevance of content in different languages. We then conduct granular ablations by reintroducing parallel or code-switching data into the monolingual-only corpus. Our experiments reveal that parallel data almost fully restores translation performance (91% of the unfiltered baseline), whereas code-switching contributes minimally. Other cross-lingual tasks remain largely unaffected by either type. These findings reveal that translation critically depends on systematic token-level alignments from parallel data, whereas cross-lingual understanding and reasoning appear to be achievable even without bilingual data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e662fab-af63-4123-bb66-221c539b8501Builds on5
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation CapabilityEleftheria Briakou, Colin Cherry, George F. FosterACL 2023 · 16 citations
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language ModelsMuhammad Reza Qorib, Junyi Li, Hwee Tou NgACL 2025 · 5 citations
Related papers
- The Reasonableness Behind Unreasonable Translation Capability of Large Language ModelTingchen Fu, Lemao Liu, Deng Cai, Guoping Huang et al.ICLR 2024
- Multilingual Language Model Pretraining using Machine-translated DataJiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin et al.EMNLP 2025
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu et al.AAAI 2020 · 94 citations
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder et al.ACL 2021
- PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual AlignmentJiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai et al.EMNLP 2024 · 5 citations
