Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
Ping Guo, Yubing Ren, Binbin Liu, Fengze Liu, Haobin Lin, Yifan Zhang, Bingni Zhang, Taifeng Wang, Yin Zheng
Abstract
Large language models (LLMs) have become integral to a wide range of applications worldwide, driving an unprecedented global demand for effective multilingual capabilities. Central to achieving robust multilingual performance is the strategic allocation of language proportions within training corpora. However, determining optimal language ratios is highly challenging due to intricate cross-lingual interactions and sensitivity to dataset scale. This paper introduces Climb (Cross-Lingual Interaction-aware Multilingual Balancing), a novel framework designed to systematically optimize multilingual data allocation. At its core, Climb introduces a cross-lingual interaction-aware language ratio, explicitly quantifying each language's effective allocation by capturing inter-language dependencies. Leveraging this ratio, Climb proposes a principled two-step optimization procedure--first equalizing marginal benefits across languages, then maximizing the magnitude of the resulting language allocation vectors--significantly simplifying the inherently complex multilingual optimization problem. Extensive experiments confirm that Climb can accurately measure cross-lingual interactions across various multilingual settings. LLMs trained with Climb-derived proportions consistently achieve state-of-the-art multilingual performance, even achieving competitive performance with open-sourced LLMs trained with more tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model PretrainingZhixun Chen, Ping Guo, Wenhan Han, Yifan Zhang et al.NeurIPS 2025 · 4 citations
- Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual TranslationBinbin Liu, Wenhan Han, Feng Chen, Yifan Zhang et al.ICLR 2026
Builds on28
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- Scaling Laws for Fine-Grained Mixture of ExpertsJan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro et al.ICML 2024 · 149 citations
Related papers
- Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchACL 2025 · 17 citations
- SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkJunpeng Liu, Jiuyi Li, Kaiyu Huang, Bo Jin et al.ACL 2026
- Language Imbalance Driven Rewarding for Multilingual Self-improvingWen Yang, Junhong Wu, Chen Wang, Chengqing Zong et al.ICLR 2025
- CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng et al.ACL 2025 · 3 citations
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentMengyu Bu, Shaolei Zhang, Zhongjun He, Hua Wu et al.EMNLP 2025
