ACL2026
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Wenhan Han, Yifan Zhang, Zhixun Chen, Binbin Liu, Mykola Pechenizkiy, Meng Fang, Yin Zheng
被引用 11 次
摘要
Multilingual large language models (LLMs) are advancing rapidly, with new models frequently claiming support for an increasing number of languages. However, existing evaluation datasets are limited and lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. To address this, we introduce MUBENCH, a benchmark covering 61 languages and evaluating a broad range of capabilities. We evaluate several state-of-the-art multilingual LLMs and find notable gaps between claimed and actual language coverage, particularly a persistent performance disparity between English and low-resource languages. Leveraging MuBench's alignment, we propose Multilingual Consistency (MLC) as a complementary metric to accuracy for analyzing performance bottlenecks and guiding model improvement. Finally, we pretrain a suite of 1.2B-parameter models on English and Chinese with 500B tokens, varying language ratios and parallel data proportions to investigate cross-lingual transfer dynamics. Our dataset will be open at https://huggingface.co/datasets/aialt/MuBench