Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLU
Fajri Koto, Nurul Aisyah, Haonan Li, Timothy Baldwin
摘要
Although large language models (LLMs) are often pre-trained on large-scale multilingual texts, their reasoning abilities and real-world knowledge are mainly evaluated based on English datasets. Assessing LLM capabilities beyond English is increasingly vital but hindered due to the lack of suitable datasets. In this work, we introduce IndoMMLU, the first multi-task language understanding benchmark for Indonesian culture and languages, which consists of questions from primary school to university entrance exams in Indonesia. By employing professional teachers, we obtain 14,981 questions across 64 tasks and education levels, with 46% of the questions focusing on assessing proficiency in the Indonesian language and knowledge of nine local languages and cultures in Indonesia. Our empirical evaluations show that GPT-3.5 only manages to pass the Indonesian primary school level, with limited knowledge of local Indonesian languages and culture. Other smaller models such as BLOOMZ and Falcon perform at even lower levels. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani 等ACL 2025 · 被引用 144 次
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh 等EMNLP 2024 · 被引用 21 次
- EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language ModelsRocktim Jyoti Das, Simeon Emilov Hristov, Haonan Li, Dimitar Dimitrov 等ACL 2024 · 被引用 13 次
- KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of KazakhstanMukhammed Togmanov, Nurdaulet Mukhituly, Diana Turmakhan, Jonibek Mansurov 等ACL 2025 · 被引用 10 次
- Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and SundaneseRifki Afina Putri, Faiz Ghifari Haznitrama, Dea Adhista, Alice OhEMNLP 2024 · 被引用 7 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMsCuc Thi Bui, Nguyen Truong Son, Trang Van Truong, Viet Lam Phung 等ACL 2025
- LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian LanguagesAlham Fikri Aji, Trevor CohnEMNLP 2025 · 被引用 2 次
- IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language GenerationSamuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio 等EMNLP 2021 · 被引用 85 次
- SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in SinhalaAshmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage 等EMNLP 2025 · 被引用 1 次
- IndoSafety: Culturally Grounded Safety for LLMs in Indonesian LanguagesMuhammad Falensi Azmi, Muhammad Dehan Al Kautsar, Alfan Farizki Wicaksono, Fajri KotoEMNLP 2025 · 被引用 1 次
