OLA: Output Language Alignment in Code-Switched LLM Interactions
Juhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice Oh
摘要
Code-switching, alternating between languages within a conversation, is natural for multilingual users, yet poses fundamental challenges for large language models (LLMs). When a user code-switches in their prompt to an LLM, they typically do not specify the expected language of the LLM response, and thus LLMs must infer the output language from contextual and pragmatic cues. We find that current LLMs systematically fail to align with this expectation, responding in undesired languages even when cues are clear to humans. We introduce OLA, a benchmark to evaluate LLMs' Output Language Alignment in codeswitched interactions. OLA focuses on Korean-English code-switching and spans simple intrasentential mixing to instruction-content mismatches. Even frontier models frequently misinterpret implicit language expectation, exhibiting a bias toward non-English responses. We further show this bias generalizes beyond Korean to Chinese and Indonesian pairs. Models also show instability through mid-response switching and language intrusions. Chain-of-Thought prompting fails to resolve these errors, indicating weak pragmatic reasoning about output language. However, Code-Switching Aware DPO with minimal data (∼1K examples) substantially reduces misalignment, suggesting these failures stem from insufficient alignment rather than fundamental limitations. Our results highlight the need to align multilingual LLMs with users' implicit expectations in real-world code-switched interactions. 1 * Equal contribution. 1 OLA is available at https://github.com/juhyunohh/ OLA replies to be in the language of the original message → Response language: English 안녕하세요. 본 운동(Main Workout) 전에 워밍업 세트(Warm-up Sets)를 수행하는 것은 부상을 방지하고 운동 효과를 [...] Translation: Hello. Performing warm-up sets before the main workout is very important for preventing injuries and maximizing [...] )
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- Do Multilingual Users Prefer Chat-bots that Code-mix? Let's Nudge and Find Out!Anshul Bawa, Pranav Khadpe, Pratik Joshi, Kalika Bali 等CSCW 2020 · 被引用 41 次
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata 等EMNLP 2023 · 被引用 19 次
相关 Paper
- Minimal Pair-Based Evaluation of Code-SwitchingIgor Sterner, Simone TeufelACL 2025 · 被引用 8 次
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMsBoyi Deng, Yu Wan, Baosong Yang, Fei Huang 等ICLR 2026 · 被引用 2 次
- Lost in the Mix: Evaluating LLM Understanding of Code-Switched TextAmr Mohamed, Yang Zhang, Michalis Vazirgiannis, Guokan ShangACL 2026 · 被引用 9 次
- CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 LanguagesYilun Yang, Yekun ChaiEMNLP 2025 · 被引用 1 次
- The Impact of Language Mixing on Bilingual LLM ReasoningYihao Li, Jiayi Xin, Miranda Muqing Miao, Qi Long 等EMNLP 2025 · 被引用 8 次
