Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
Collin Zhang, Fei Huang, Chenhan Yuan, Junyang Lin
Abstract
Large language models (LLMs) often experience language confusion, which is the unintended mixing of languages during text generation. Current solutions to this problem either necessitate model retraining or cannot differentiate between harmful confusion and acceptable code-switching. This paper introduces the Language Confusion Gate (LCG), a lightweight, plug-in solution that filters tokens during decoding without altering the base LLM. The LCG is trained using norm-adjusted self-distillation to predict appropriate language families and apply masking only when needed. Our method is based on the findings that language confusion is infrequent, correct-language tokens are usually among the top predictions, and output token embedding norms are larger for high-resource languages, which biases sampling. When evaluated across various models, including Qwen3, GPT-OSS, Gemma3, Llama3.1, LCG decreases language confusion significantly, often by an order of magnitude, without negatively impacting task performance. Code is available at https://github.com/collinzrj/language_confusion_gate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e46a1ae-2c15-4b05-b06e-7bc6556c8badBuilds on4
- Understanding and Mitigating Language Confusion in LLMsKelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze et al.EMNLP 2024 · 10 citations
- The Impact of Language Mixing on Bilingual LLM ReasoningYihao Li, Jiayi Xin, Miranda Muqing Miao, Qi Long et al.EMNLP 2025 · 8 citations
- A Survey of Code-switching: Linguistic and Social Perspectives for Language TechnologiesA. Seza Dogruöz, Sunayana Sitaram, Barbara E. Bullock, Almeida Jacqueline ToribioACL 2021
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningShivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson et al.ACL 2024
Related papers
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMsBoyi Deng, Yu Wan, Baosong Yang, Fei Huang et al.ICLR 2026 · 2 citations
- TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language ModelsJinho Choo, JunSeung Lee, Jimyeong Kim, Yeeho Song et al.ACL 2026
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 LanguagesYilun Yang, Yekun ChaiEMNLP 2025 · 1 citation
- Quantification of Large Language Model DistillationSunbowen Lee, Junting Zhou, Chang Ao, Kaige Li et al.ACL 2025
- Token-Level LLM Collaboration via FusionRouteNuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen et al.ICML 2026 · 7 citations
