MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models
Yan Cai, Linlin Wang, Ye Wang, Gerard de Melo, Ya Zhang, Yanfeng Wang, Liang He
Abstract
The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the Chinese medical domain, comprising 40,041 questions sourced from authentic examination exercises and medical reports of diverse branches of medicine. In particular, this benchmark is composed of four key components: the Chinese Medical Licensing Examination, the Resident Standardization Training Examination, the Doctor In-Charge Qualification Examination, and real-world clinic cases encompassing examinations, diagnoses, and treatments. MedBench replicates the educational progression and clinical practice experiences of doctors in Mainland China, thereby establish- ing itself as a credible benchmark for assessing the mastery of knowledge and reasoning abilities in medical language learning models. We perform extensive experiments and conduct an in-depth analysis from diverse perspectives, which culminate in the following findings: (1) Chinese medical LLMs underperform on this benchmark, highlighting the need for significant advances in clinical knowledge and diagnostic precision. (2) Several general-domain LLMs surprisingly possess considerable medical knowledge. These findings elucidate both the capabilities and limitations of LLMs within the context of MedBench, with the ultimate goal of aiding the medical research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e2dd64e-d511-4da9-9de1-bfaf31e44dfdCited by top-tier papers8
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo et al.EMNLP 2024 · 6 citations
- Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical EvaluationShunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang et al.AAAI 2025 · 4 citations
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic EvaluationXiangxu Zhang, Lei Li, Yanyun Zhou, Xiao Zhou et al.ACL 2026 · 3 citations
- Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question AnsweringXueren Ge, Sahil Murtaza, Anthony Cortez, Homa AlemzadehAAAI 2026 · 2 citations
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning et al.ICML 2025
Builds on4
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 112 citations
- MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering DatasetJing Li, Shangping Zhong, Kaizhi ChenEMNLP 2021 · 24 citations
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
Related papers
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 1 citation
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- CBLUE: A Chinese Biomedical Language Understanding Evaluation BenchmarkNingyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang et al.ACL 2022 · 242 citations
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai et al.ACL 2025
