Respond in my Language: Mitigating Language Inconsistency in Response Generation based on Large Language Models
Liang Zhang, Qin Jin, Haoyang Huang, Dongdong Zhang, Furu Wei
Abstract
Large Language Models (LLMs) show strong instruction understanding ability across multiple languages. However, they are easily biased towards English in instruction tuning, and generate English responses even given non-English instructions. In this paper, we investigate the language inconsistent generation problem in monolingual instruction tuning. We find that instruction tuning in English increases the models' preference for English responses. It attaches higher probabilities to English responses than to responses in the same language as the instruction. Based on the findings, we alleviate the language inconsistent generation problem by counteracting the model preference for English responses in both the training and inference stages. Specifically, we propose Pseudo-Inconsistent Penalization (PIP) which prevents the model from generating English responses when given non-English language prompts during training, and Prior Enhanced Decoding (PED) which improves the language-consistent prior by leveraging the untuned base language model. Experimental results show that our two methods significantly improve the language consistency of the model without requiring any multilingual data 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive FrameworkHengyuan Zhang, Chenming Shang, Sizhe Wang, Dongdong Zhang et al.ACL 2025 · 5 citations
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- Parrot: Multilingual Visual Instruction TuningHai-Long Sun, Da-Wei Zhou, Yang Li, Shiyin Lu et al.ICML 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
Related papers
- PLUG: Leveraging Pivot Language in Cross-Lingual Instruction TuningZhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu et al.ACL 2024
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- Understanding and Mitigating Language Confusion in LLMsKelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze et al.EMNLP 2024 · 10 citations
- LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint GuidanceYuchun Fan, Bei Li, Peiguang Li, Yilin Wang et al.ACL 2026 · 1 citation
- SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkJunpeng Liu, Jiuyi Li, Kaiyu Huang, Bo Jin et al.ACL 2026
