ACL2026

BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical Scores

Jiajia Li, Weizhi Xue, Yao Yao, Qiwei Li, Chong Chen, Zuchao Li, Ping Wang, Hai Zhao

摘要

Multimodal Large Language Models (MLLMs) excel in general tasks but struggle with specialized, structured cultural symbols. We introduce BoYaEval, the first comprehensive benchmark dedicated to deciphering diverse Ancient Chinese musical notations, including five types of ancient Chinese music notation systems. These systems utilize unique spatial layouts and specialized ideograms to encode pitch and intricate playing techniques. BoYaEval comprises 3,175 high-quality images across these notation styles and establishes a three-tier evaluation: Structural Parsing (symbol recognition), Instructional Translation (technique mapping), and Musical Reasoning (melody derivation). We evaluate 21 leading MLLMs. Results indicate that while models perform adequately in basic recognition, they fail in cross-system compositional logic, scoring only around 27% on reasoning tasks. BoYaEval highlights the limitations of current MLLMs in processing diverse spatial-symbolic dependencies, bridging the gap between ancient wisdom and modern AI for digitizing intangible cultural heritage. The BoYaEval benchmark is publicly available at https://huggingface. co/datasets/MYTH-Lab/BoYaEval . Jianzipu Jianzipu uses composite ideograms. Each "glyph" is a combination of fragments taken from standard Chinese characters. These fragments are shrunk together to form a single block that functions as a performance instruction.