Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark
Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag
Abstract
The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the closeended question-answering (QA) task with answer options for evaluation. However, many clinical decisions involve answering openended questions without pre-set options. To better understand LLMs in the clinic, we construct a benchmark ClinicBench. We first collect eleven existing datasets covering diverse clinical language generation, understanding, and reasoning tasks. Furthermore, we construct six novel datasets and clinical tasks that are complex but common in real-world practice, e.g., open-ended decision-making, long document processing, and emerging drug analysis. We conduct an extensive evaluation of twentytwo LLMs under both zero-shot and few-shot settings. Finally, we invite medical experts to evaluate the clinical usefulness of LLMs 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health RecordsDongchen Li, Jitao Liang, Wei Li, Xiaoyu Wang et al.AAAI 2026 · 1 citation
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks AlignmentJean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha et al.ACL 2025
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-World Multi-Turn DialogueSonghua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou et al.AAAI 2024 · 227 citations
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 112 citations
Related papers
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo et al.EMNLP 2024 · 6 citations
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo et al.AAAI 2024 · 42 citations
- CLINIC : Evaluating Multilingual Trustworthiness in Language Models for HealthcareAkash Ghosh, Srivarshinee Sridhar, Raghav Kaushik Ravi, Muhsin Muhsin et al.ICML 2026 · 6 citations
- Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in MedicineMaxime Griot, Jean Vanderdonckt, Demet Yüksel, Coralie HemptinneACL 2025
