Understanding and Mitigating Language Confusion in LLMs
Kelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze, Sebastian Ruder
Abstract
We investigate a surprising limitation of LLMs: their inability to consistently generate text in a user's desired language. We create the Language Confusion Benchmark (LCB) to evaluate such failures, covering 15 typologically diverse languages with existing and newly-created English and multilingual prompts. We evaluate a range of LLMs on monolingual and crosslingual generation reflecting practical use cases, finding that Llama Instruct and Mistral models exhibit high degrees of language confusion and even the strongest models fail to consistently respond in the correct language. We observe that base and English-centric instruct models are more prone to language confusion, which is aggravated by complex prompts and high sampling temperatures. We find that language confusion can be partially mitigated via fewshot prompting, multilingual SFT and preference tuning. We release our language confusion benchmark, which serves as a first layer of efficient, scalable multilingual evaluation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context LearningErcong Nie, Bo Shao, Mingyang Wang, Zifeng Ding et al.ACL 2025 · 14 citations
- The Impact of Language Mixing on Bilingual LLM ReasoningYihao Li, Jiayi Xin, Miranda Muqing Miao, Qi Long et al.EMNLP 2025 · 8 citations
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 4 citations
- Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time MarkersDaniel D'souza, Julia Kreutzer, Adrien Morisot, Ahmet Üstün et al.NeurIPS 2025 · 3 citations
- MENLO: From Preferences to Proficiency - Evaluating and Modeling Native-like Quality Across 47 LanguagesChenxi Whitehouse, Sebastian Ruder, Tony Lin, Oksana Kurylo et al.ICLR 2026 · 3 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye et al.ICML 2024 · 274 citations
Related papers
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu et al.ACL 2025 · 5 citations
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in InstructionsXingwei He, Qianru Zhang, Pengfei Chen, Guanhua Chen et al.AAAI 2026 · 2 citations
- Multi-LCB: Extending LiveCodeBench to Multiple Programming LanguagesMaria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov et al.ICLR 2026 · 3 citations
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng et al.ACL 2025 · 3 citations
- LLMs can be easily Confused by Instructional DistractionsYerin Hwang, Yongil Kim, Jahyun Koo, Taegwan Kang et al.ACL 2025
