How do Large Language Models Handle Multilingualism?
Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, Lidong Bing
摘要
Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow (): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify , we introduce Parallel Language-specific Neuron Detection () to identify activated neurons for inputs in different languages without any labeled data. Using , we validate through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of for high-resource languages and for low-resource languages across all tasks with just documents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersWeixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu 等NeurIPS 2025 · 被引用 20 次
- The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning ModelsRenfei Dang, Zhening Li, Shujian Huang, Jiajun ChenICLR 2026 · 被引用 10 次
- Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety AlignmentYuyan Bu, Xiaohao Liu, ZhaoXing Ren, Yaodong Yang 等ICLR 2026 · 被引用 9 次
- The Impact of Language Mixing on Bilingual LLM ReasoningYihao Li, Jiayi Xin, Miranda Muqing Miao, Qi Long 等EMNLP 2025 · 被引用 8 次
- Exploring the Translation Mechanism of Large Language ModelsHongbin Zhang, Kehai Chen, Xuefeng Bai, Xiucheng Li 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper17
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank ReductionPratyusha Sharma, Jordan T. Ash, Dipendra MisraICLR 2024 · 被引用 135 次
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 被引用 92 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
相关 Paper
- Revealing the Parallel Multilingual Learning within Large Language ModelsYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu 等EMNLP 2024
- The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageYuxin Chen, Yiran Zhao, Yang Zhang, An Zhang 等NeurIPS 2025 · 被引用 27 次
- Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsTianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang 等ACL 2024
- 1+12: Can Large Language Models Serve as Cross-Lingual Knowledge Aggregators?Yue Huang, Chenrui Fan, Yuan Li, Siyuan Wu 等EMNLP 2024 · 被引用 1 次
- How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons PerspectiveShimao Zhang, Zhejian Lai, Xiang Liu, Shuaijie She 等AAAI 2026 · 被引用 4 次
