MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical Studies
Yang Liu, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi, Kai Ding, Lianwen Jin
摘要
With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Chinese Classical Studies (CCS), a field which plays a vital role in preserving and promoting China's rich cultural heritage, remains largely unexplored due to the absence of specialized benchmarks. To bridge this gap, we propose MCS-Bench, the first-of-its-kind multimodal benchmark specifically designed for CCS across multiple subdomains. MCS-Bench spans seven core subdomains (Ancient Chinese Text, Calligraphy, Painting, Oracle Bone Script, Seal, Cultural Relic, and Illustration), with a total of 45 meticulously designed tasks. Through extensive evaluation of 37 representative MLLMs, we observe that even the topperforming model (InternVL2.5-78B) achieves an average score below 50, indicating substantial room for improvement. Our analysis reveals significant performance variations across different tasks and identifies critical challenges in areas such as Optical Character Recognition (OCR) and cultural context interpretation. MCS-Bench not only establishes a standardized baseline for CCS-focused MLLM research but also provides valuable insights for advancing cultural heritage preservation and innovation in the Artificial General Intelligence (AGI) era. Data and code will be publicly available. Question Fromat Method Dataset Domain Modality License Scale # Category # Task # LLM MCQ QA HG CI MC C-Eval General Text-only CC BY-NC-SA-4.0 439 1 2 11 Chinese SimpleQA General Text-only CC BY-NC-SA-4.0 323 4 11 41 CIF-Bench General Text-only -150 1 3 28 CMMLU General Text-only CC BY-NC-4.0 1,192 1 7 21 GAOKAO-Bench General Text-only Apache-2.0 145 1 2 12 XiezhiBenchmark General Text-only CC BY-NC-SA-4.0 2,060 2 3 47 ACLUE CCS Text-only CC BY-NC-4.0 4,967 5 15 8 C-CLUE CCS Text-only CC BY-SA-4.0 1,122 1 2 -CCLUE CCS Text-only Apache-2.0 36,319 2 5 -CCPM CCS Text-only -2,720 1 1 -THUAIPoet CCS Text-only -5,173 1 3 -WenMind CCS Text-only CC BY-NC-SA-4.0 4,875 3 42 31 WYWEB CCS Text-only -69,700 5 9 -CII-Bench General Image-Text Apache-2.0 137 1 1 21 ALM-Bench Culture Image-Text CC BY-NC-4.0 466 18 -16 CulturalBench Culture Image-Text CC BY-4.0 117 3 -18 CVQA Culture Image-Text -311 10 -16 MaRVL Culture Image-Text CC BY-4.0 1,012 11 1 -MCS-Bench (Ours) CCS Image-Text CC BY-NC-SA-4.0 6,500 7 45 37
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li 等ACL 2026 · 被引用 1 次
- MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical DocumentsYijun Sheng, Shipeng Zhu, Ruijia Zuo, Na Nie 等CVPR 2026
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li 等ICCV 2025 · 被引用 37 次
- CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language ModelWei Zhang, Wong Kam-Kwai, Biying Xu, Yiwen Ren 等ACM MM 2025 · 被引用 3 次
相关 Paper
- Can MLLMs Understand the Deep Implication Behind Chinese Images?Chenhao Zhang, Xi Feng, Yuelin Bai, Xeron Du 等ACL 2025
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi 等AAAI 2026 · 被引用 1 次
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li 等ACL 2026
- TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningJiahuan Cao, Yang Liu, Peirong Zhang, Yongxin Shi 等ACM MM 2025
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?Zijian Chen, Tingzhu Chen, Wenjun Zhang, Guangtao ZhaiICLR 2025 · 被引用 3 次
