Language of Thought Shapes Output Diversity in Large Language Models
Shaoyang Xu, Wenxuan Zhang
Abstract
Output diversity is crucial for Large Language Models as it underpins pluralism and creativity. In this work, we reveal that controlling the language used during model thinking-the language of thought-provides a novel and structural source of output diversity. Our preliminary study shows that different thinking languages occupy distinct regions in a model's thinking space. Based on this observation, we study two repeated sampling strategies under multilingual thinking-Single-Language Sampling and Mixed-Language Sampling-and conduct diversity evaluation on outputs that are controlled to be in English, regardless of the thinking language used. Across extensive experiments, we demonstrate that switching the thinking language from English to non-English languages consistently increases output diversity, with a clear and consistent positive correlation such that languages farther from English in the thinking space yield larger gains. We further show that aggregating samples across multiple thinking languages yields additional improvements through compositional effects, and that scaling sampling with linguistic heterogeneity expands the model's diversity ceiling. Finally, we show that these findings translate into practical benefits in pluralistic alignment scenarios, leading to broader coverage of cultural knowledge and value orientations in LLM outputs. Our code is publicly available at https://github.com/ iNLP-Lab/Multilingual-LoT-Diversity .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
- Does Writing with Language Models Reduce Content Diversity?Vishakh Padmakumar, He HeICLR 2024 · 173 citations
- Linguistic Generalizability of Test-Time Scaling in Mathematical ReasoningGuijin Son, Jiwoo Hong, Hyunwoo Ko, James ThorneACL 2025 · 36 citations
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li et al.EMNLP 2025 · 33 citations
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
Related papers
- Generative Monoculture in Large Language ModelsFan Wu, Emily Black, Varun ChandrasekaranICLR 2025
- OLA: Output Language Alignment in Code-Switched LLM InteractionsJuhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice OhACL 2026 · 1 citation
- Do Llamas Work in English? On the Latent Language of Multilingual TransformersChris Wendler, Veniamin Veselovsky, Giovanni Monea, Robert WestACL 2024
- Multilingual Jailbreak Challenges in Large Language ModelsYue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong BingICLR 2024 · 230 citations
- Revealing the Parallel Multilingual Learning within Large Language ModelsYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu et al.EMNLP 2024
