MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical Studies
Yang Liu, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi, Kai Ding, Lianwen Jin
Abstract
With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Chinese Classical Studies (CCS), a field which plays a vital role in preserving and promoting China's rich cultural heritage, remains largely unexplored due to the absence of specialized benchmarks. To bridge this gap, we propose MCS-Bench, the first-of-its-kind multimodal benchmark specifically designed for CCS across multiple subdomains. MCS-Bench spans seven core subdomains (Ancient Chinese Text, Calligraphy, Painting, Oracle Bone Script, Seal, Cultural Relic, and Illustration), with a total of 45 meticulously designed tasks. Through extensive evaluation of 37 representative MLLMs, we observe that even the topperforming model (InternVL2.5-78B) achieves an average score below 50, indicating substantial room for improvement. Our analysis reveals significant performance variations across different tasks and identifies critical challenges in areas such as Optical Character Recognition (OCR) and cultural context interpretation. MCS-Bench not only establishes a standardized baseline for CCS-focused MLLM research but also provides valuable insights for advancing cultural heritage preservation and innovation in the Artificial General Intelligence (AGI) era. Data and code will be publicly available. Question Fromat Method Dataset Domain Modality License Scale # Category # Task # LLM MCQ QA HG CI MC C-Eval General Text-only CC BY-NC-SA-4.0 439 1 2 11 Chinese SimpleQA General Text-only CC BY-NC-SA-4.0 323 4 11 41 CIF-Bench General Text-only -150 1 3 28 CMMLU General Text-only CC BY-NC-4.0 1,192 1 7 21 GAOKAO-Bench General Text-only Apache-2.0 145 1 2 12 XiezhiBenchmark General Text-only CC BY-NC-SA-4.0 2,060 2 3 47 ACLUE CCS Text-only CC BY-NC-4.0 4,967 5 15 8 C-CLUE CCS Text-only CC BY-SA-4.0 1,122 1 2 -CCLUE CCS Text-only Apache-2.0 36,319 2 5 -CCPM CCS Text-only -2,720 1 1 -THUAIPoet CCS Text-only -5,173 1 3 -WenMind CCS Text-only CC BY-NC-SA-4.0 4,875 3 42 31 WYWEB CCS Text-only -69,700 5 9 -CII-Bench General Image-Text Apache-2.0 137 1 1 21 ALM-Bench Culture Image-Text CC BY-NC-4.0 466 18 -16 CulturalBench Culture Image-Text CC BY-4.0 117 3 -18 CVQA Culture Image-Text -311 10 -16 MaRVL Culture Image-Text CC BY-4.0 1,012 11 1 -MCS-Bench (Ours) CCS Image-Text CC BY-NC-SA-4.0 6,500 7 45 37
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41a9b92f-978a-4c09-8adb-76e1e2588067Cited by top-tier papers2
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li et al.ACL 2026 · 1 citation
- MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical DocumentsYijun Sheng, Shipeng Zhu, Ruijia Zuo, Na Nie et al.CVPR 2026
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy et al.EMNLP 2021 · 87 citations
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li et al.ICCV 2025 · 37 citations
- CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language ModelWei Zhang, Wong Kam-Kwai, Biying Xu, Yiwen Ren et al.ACM MM 2025 · 3 citations
Related papers
- Can MLLMs Understand the Deep Implication Behind Chinese Images?Chenhao Zhang, Xi Feng, Yuelin Bai, Xeron Du et al.ACL 2025
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi et al.AAAI 2026 · 1 citation
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li et al.ACL 2026
- TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningJiahuan Cao, Yang Liu, Peirong Zhang, Yongxin Shi et al.ACM MM 2025
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?Zijian Chen, Tingzhu Chen, Wenjun Zhang, Guangtao ZhaiICLR 2025 · 3 citations
