The Impact of Language Mixing on Bilingual LLM Reasoning
Yihao Li, Jiayi Xin, Miranda Muqing Miao, Qi Long, Lyle H. Ungar
Abstract
Proficient multilingual speakers often intentionally switch languages in the middle of a conversation. Similarly, recent reasoning-focused bilingual large language models (LLMs) with strong capabilities in both languages exhibit language mixing-alternating languages within their chain of thought. Discouraging this behavior in DeepSeek-R1 was found to degrade accuracy, suggesting that language mixing may benefit reasoning. In this work, we study language switching in Chinese-English bilingual reasoning models. We identify reinforcement learning with verifiable rewards (RLVR) as the critical training stage that leads to language mixing. We show that language mixing can enhance reasoning: enforcing monolingual decoding reduces accuracy by 5.6 percentage points on MATH500. Additionally, a lightweight probe can be trained to predict whether a potential language switch would benefit or harm reasoning, and when used to guide decoding, increases accuracy by 2.92 percentage points. Our findings suggest that language mixing is not merely a byproduct of multilingual training, but is a strategic reasoning behavior. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e0301f7-005f-4516-9b25-a0b0fe0445daCited by top-tier papers4
- Language of Thought Shapes Output Diversity in Large Language ModelsShaoyang Xu, Wenxuan ZhangACL 2026 · 3 citations
- OLA: Output Language Alignment in Code-Switched LLM InteractionsJuhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice OhACL 2026 · 1 citation
- Language Confusion Gate: Language-Aware Decoding Through Model Self-DistillationCollin Zhang, Fei Huang, Chenhan Yuan, Junyang LinICLR 2026 · 1 citation
- Learning to Route Languages for Multilingual Policy OptimizationGeyang Guo, Hiromi Wakaki, Yuki Mitsufuji, Alan Ritter et al.ICML 2026
Builds on8
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi et al.NeurIPS 2024 · 196 citations
- Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal ProofsAlbert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, Timothée Lacroix et al.ICLR 2023 · 25 citations
- Understanding and Mitigating Language Confusion in LLMsKelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze et al.EMNLP 2024 · 10 citations
- Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-trainingLinjuan Wu, Haoran Wei, Huan Lin, Tianhao Li et al.EMNLP 2025 · 1 citation
Related papers
- Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal CausesMingyang Wang, Lukas Lange, Heike Adel, Yunpu Ma et al.EMNLP 2025
- LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint GuidanceYuchun Fan, Bei Li, Peiguang Li, Yilin Wang et al.ACL 2026 · 1 citation
- When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersWeixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu et al.NeurIPS 2025 · 20 citations
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling et al.ICLR 2026 · 112 citations
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
