CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni
Abstract
Artificial Intelligence (AI), along with the recent progress in biomedical language understanding, is gradually offering great promise for medical practice. With the development of biomedical language understanding benchmarks, AI applications are widely used in the medical field. However, most benchmarks are limited to English, which makes it challenging to replicate many of the successes in English for other languages. To facilitate research in this direction, we collect real-world biomedical data and present the first Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark: a collection of natural language understanding tasks including named entity recognition, information extraction, clinical diagnosis normalization, single-sentence/sentence-pair classification, and an associated online platform for model evaluation, comparison, and analysis. To establish evaluation on these tasks, we report empirical results with the current 11 pre-trained Chinese models, and experimental results show that state-of-the-art neural models perform by far worse than the human ceiling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b9bf3f7-594f-4058-ba94-c858ebf5efc1Cited by top-tier papers12
- D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented ChatBinwei Yao, Chao Shi, Likai Zou, Lingfeng Dai et al.EMNLP 2022 · 23 citations
- Multitask Pre-training of Modular Prompt for Chinese Few-Shot LearningTianxiang Sun, Zhengfu He, Qin Zhu, Xipeng Qiu et al.ACL 2023 · 15 citations
- MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity RecognitionJinyuan Fang, Xiaobin Wang, Zaiqiao Meng, Pengjun Xie et al.ACL 2023 · 13 citations
- Manifold-Based Verbalizer Space Re-embedding for Tuning-Free Prompt-Based ClassificationHaochun Wang, Sendong Zhao, Chi Liu, Nuwa Xi et al.AAAI 2024 · 4 citations
- Generative Models for Automatic Medical Decision Rule Extraction from TextYuxin He, Buzhou Tang, Xiaoling WangEMNLP 2024 · 2 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Graph-Evolving Meta-Learning for Low-Resource Medical Dialogue GenerationShuai Lin, Pan Zhou, Xiaodan Liang, Jianheng Tang et al.AAAI 2021 · 66 citations
- RussianSuperGLUE: A Russian Language Understanding Evaluation BenchmarkTatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev et al.EMNLP 2020 · 11 citations
Related papers
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo et al.AAAI 2024 · 42 citations
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 1 citation
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo et al.EMNLP 2024 · 6 citations
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang et al.AAAI 2025 · 7 citations
- A Knowledge-driven Generative Model for Multi-implication Chinese Medical Procedure Entity NormalizationJinghui Yan, Yining Wang, Lu Xiang, Yu Zhou et al.EMNLP 2020 · 17 citations
