The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model
Jiawei Chen, Wentao Chen, Jing Su, Jingjing Xu, Hongyu Lin, Mengjie Ren, Yaojie Lu, Xianpei Han, Le Sun
摘要
Large language models (LLMs) have shown significant multilingual capabilities. However, the mechanisms underlying the development of these capabilities during pre-training are not well understood. In this paper, we use code LLMs as an experimental platform to explore the evolution of multilingual capabilities in LLMs during the pre-training process. Based on our observations, we propose the Babel Tower Hypothesis, which describes the entire process of LLMs acquiring new language capabilities. During the learning process, multiple languages initially share a single knowledge system dominated by the primary language and gradually develop language-specific knowledge systems. We then validate the above hypothesis by tracking the internal states of the LLMs through identifying working languages and language transferring neurons. Experimental results show that the internal state changes of the LLM are consistent with our Babel Tower Hypothesis. Building on these insights, we propose a novel method to construct an optimized pre-training corpus for multilingual code LLMs, which significantly outperforms LLMs trained on the original corpus. The proposed Babel Tower Hypothesis provides new insights into designing pre-training data distributions to achieve optimal multilingual capabilities in LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsXiang Zhang, Senyu Li, Bradley Hauer, Ning Shi 等EMNLP 2023 · 被引用 58 次
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li 等ICLR 2023 · 被引用 28 次
- Multilingual LLMs are Better Cross-lingual In-context Learners with AlignmentEshaan Tanwar, Subhabrata Dutta, Manish Borthakur, Tanmoy ChakrabortyACL 2023 · 被引用 21 次
- Cross-Lingual Consistency of Factual Knowledge in Multilingual Language ModelsJirui Qi, Raquel Fernández, Arianna BisazzaEMNLP 2023 · 被引用 9 次
- Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A ReviewFred Philippy, Siwen Guo, Shohreh HaddadanACL 2023 · 被引用 9 次
相关 Paper
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi 等NeurIPS 2024 · 被引用 196 次
- Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsTianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang 等ACL 2024
- Cross-Lingual Generalization and Compression: From Language-Specific to Shared NeuronsFrederick Riemenschneider, Anette FrankACL 2025 · 被引用 3 次
- One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersDiana Abagyan, Alejandro Salamanca, Andrés Felipe Cruz-Salinas, Kris Cao 等ACL 2026 · 被引用 11 次
- MTLS: Making Texts into Linguistic SymbolsWenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang 等EMNLP 2024 · 被引用 1 次
