Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre
摘要
Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also from the cultural knowledge required to interpret questions, reducing the practical utility of translated datasets like MMLU. Furthermore, translation often introduces artefacts that can distort the meaning or clarity of questions in the target language. A common practice in multilingual evaluation is to rely on machine-translated evaluation sets, but simply translating a dataset is insufficient to address these challenges. In this work, we trace the impact of both of these issues on multilingual evaluations and ensuing model performances. Our large-scale evaluation of state-of-the-art open and proprietary models illustrates that progress on MMLU depends heavily on learning Western-centric concepts, with 28% of all questions requiring culturally sensitive knowledge. Moreover, for questions requiring geographic knowledge, an astounding 84.9% focus on either North American or European regions. Rankings of model evaluations change depending on whether they are evaluated on the full portion or the subset of questions annotated as culturally sensitive, showing the distortion to model rankings when blindly relying on translated MMLU. We release Global MMLU, an improved MMLU with evaluation coverage across 42 languages -- with improved overall quality by engaging with compensated professional and community annotators to verify translation quality while also rigorously evaluating cultural biases present in the original dataset. This comprehensive Global MMLU set also includes designated subsets labeled as culturally sensitive and culturally agnostic to allow for more holistic, complete evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou 等ACL 2026 · 被引用 51 次
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu 等ICLR 2026 · 被引用 34 次
- When AI Benchmarks Plateau: A Systematic Study of Benchmark SaturationMubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja 等ICML 2026 · 被引用 22 次
- Pre-Trained Policy Discriminators are General Reward ModelsShihan Dou, Shichun Liu, Yuming Yang, Yicheng Zou 等NeurIPS 2025 · 被引用 13 次
- TraceRouter: Robust Safety for Large Foundation Models via Path-Level InterventionChuancheng Shi, shangze li, Wenjun Lu, Wenhua Wu 等ICML 2026 · 被引用 12 次
它引用的顶会 Paper18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 被引用 27 次
相关 Paper
- XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question AnsweringKeon-Woo Roh, Yeong-Joon Ju, Seong-Whan LeeEMNLP 2025
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu 等ACL 2026
- Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li 等ACL 2026 · 被引用 2 次
- Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual KnowledgeEshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak 等EMNLP 2025 · 被引用 4 次
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li 等EMNLP 2022 · 被引用 27 次
