Disentangling Language and Culture for Evaluating Multilingual Large Language Models
Jiahao Ying, Wei Tang, Yiran Zhao, Yixin Cao, Yu Rong, Wenxuan Zhang
摘要
This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framework enables a nuanced analysis of LLMs' ability to process questions within both native and cross-cultural contexts cross-lingually. Extensive evaluations are conducted on a wide range of models, revealing a notable "Cultural-Linguistic Synergy" phenomenon, where models exhibit better performance when questions are culturally aligned with the language. This phenomenon is further explored through interpretability probing, which shows that a higher proportion of specific neurons are activated in a language's cultural context. This activation proportion could serve as a potential indicator for evaluating multilingual performance during model training. Our findings challenge the prevailing notion that LLMs, primarily trained on English data, perform uniformly across languages and highlight the necessity of culturally and linguistically model evaluations. Our code can be found at https://yingjiahao14. github.io/Dual-Evaluation/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Language of Thought Shapes Output Diversity in Large Language ModelsShaoyang Xu, Wenxuan ZhangACL 2026 · 被引用 3 次
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot SettingsHamna, Gayatri Bhat, Sourabrata Mukherjee, Faisal M. Lalani 等CHI 2026 · 被引用 1 次
- Neuron-Level Analysis of Cultural Understanding in Large Language ModelsTaisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi YanakaICLR 2026 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- Explainability and Interpretability of Multilingual Large Language Models: A SurveyLucas Resck, Isabelle Augenstein, Anna KorhonenEMNLP 2025
- Revealing the Parallel Multilingual Learning within Large Language ModelsYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu 等EMNLP 2024
- The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao 等ACL 2026
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi 等NeurIPS 2024 · 被引用 196 次
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu 等ACL 2025
