ACL2026
MonCulture-Eval: A Hierarchical Benchmark for Evaluating Mongolian Cultural Capabilities of Large Language Models across Scripts and Regions
Quulgan Minggad, Zinan Xiao, Yuan Sun
Abstract
While Large Language Models (LLMs) have achieved impressive linguistic fluency in lowresource languages, their capacity to process deep cultural nuances remains insufficiently quantified. This paper introduces MonCulture-Eval, a benchmark designed to assess the cultural intelligence of LLMs in the Mongolian context across two writing systems (Traditional and Cyrillic) and three regional subcultures (Alxa, Ordos, and Horqin). Curated entirely from primary, non-digitized archives to inherently prevent data contamination, the benchmark employs a three-layer cognitive hierarchy-Factual, Situational, and Valuessupplemented by specialized tasks including Riddles, Taboos, and Proverbs. Evaluation of frontier models, including GPT-5.2, Gemini-3-Pro-Preview, Claude-Sonnet-4-5, DeepSeek-v3.2, and the regionally optimized Qwen3-Max, reveals distinct structural limitations. First, we observe a severe "Script Gap," where most models experience a sharp performance decline in the Traditional script, effectively restricting their access to historical cultural archives. Second, qualitative analysis identifies a systematic "Tourist Perspective" (Etic Bias), wherein models sanitize spiritual rituals into secular functional norms. While Gemini-3-Pro-preview maintains robust cross-script alignment and Emic consistency, the broader results demonstrate that linguistic translation capability does not guarantee cultural value alignment. These findings provide empirical baselines for advancing culturally grounded AI systems. Our data and code are publicly available at https:// github.com/Ayakades/MonCulture-Eval .