How Do Language Models Speak Languages? A Case Study on Unintended Code-Switching
Yuxin Xiao, Zhen Huang, Wenxiao Wang, Yan Zhao, Zhihong Gu, Binbin Lin, Xiaofei He, Xu Shen, Jieping Ye
Abstract
Unintended code-switching, where LLMs unexpectedly switch languages, poses a fundamental challenge to multilingual generation in LLMs. However, we still lack a mechanistic account of how this failure mode is implemented inside the model. Key questions remain: what internal components (i.e., circuits) give rise to unintended code-switching, where do they emerge across layers, and how can we intervene to mitigate it? In this work, we introduce a scalable circuit discovery framework that causally localizes multilingual neurons and describes their functional patterns, then further groups them into interpretable circuits—without any additional training or manual annotation. Our findings are twofold: a) The model's "speaking-a-language" circuit decomposes into a language regime (detecting and maintaining language identity) and a semantic regime (retrieving language-agnostic semantics). b) The mechanism of unintended code-switching is a regime shift. The semantic regime suppresses the language regime and overwhelms the multilingual circuit, causing the model to generate in an unintended language. To validate these findings, we further fine-tune the identified language sub-circuit, reducing the code-switching rate by with minimal parameter updates ( % of all neurons). This work serves as a preliminary exploration of multilingual generation mechanism, offering actionable insight for targeted training for multilingual LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on31
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
Related papers
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMsBoyi Deng, Yu Wan, Baosong Yang, Fei Huang et al.ICLR 2026 · 2 citations
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov et al.ICLR 2025
- LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine TranslationShaolin Zhu, Leiyu Pan, Bo Li, Deyi XiongACL 2024
- CLUE: Conflict-guided Localization for LLM Unlearning FrameworkHang Chen, Jiaying Zhu, Xinyu Yang, Wenya WangICLR 2026 · 8 citations
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata et al.EMNLP 2023 · 19 citations
