Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine
Maxime Griot, Jean Vanderdonckt, Demet Yüksel, Coralie Hemptinne
摘要
Large Language Models (LLMs) such as Chat-GPT demonstrate significant potential in the medical domain and are often evaluated using multiple-choice questions (MCQs) modeled on exams like the USMLE. However, such benchmarks may overestimate true clinical understanding by rewarding pattern recognition and test-taking heuristics. To investigate this, we created a fictional medical benchmark centered on an imaginary organ, the Glianorex, allowing us to separate memorized knowledge from reasoning ability. We generated textbooks and MCQs in English and French using leading LLMs, then evaluated proprietary, open-source, and domain-specific models in a zero-shot setting. Despite the fictional content, models achieved an average score of 64%, while physicians scored only 27%. Fine-tuned medical models outperformed base models in English but not in French. Ablation and interpretability analyses revealed that models frequently relied on shallow cues, test-taking strategies, and hallucinated reasoning to identify the correct choice. These results suggest that standard MCQ-based evaluations may not effectively measure clinical reasoning and highlight the need for more robust, clinically meaningful assessment methods for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental HealthcareMax Lamparth, Declan Grabb, Amy Franks, Scott Gershan 等ICLR 2026 · 被引用 5 次
- MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMsZhan Qu, Michael FärberACL 2026 · 被引用 1 次
它引用的顶会 Paper5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?Daniel P. Jeong, Saurabh Garg, Zachary C. Lipton, Michael OberstEMNLP 2024 · 被引用 20 次
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?Nishant Balepur, Abhilasha Ravichander, Rachel RudingerACL 2024 · 被引用 5 次
相关 Paper
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo 等AAAI 2024 · 被引用 42 次
- Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive BenchmarkFenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin 等EMNLP 2024 · 被引用 9 次
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo 等EMNLP 2024 · 被引用 6 次
- "Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Clinical ReasoningYuansong Xu, Yichao Zhu, Haokai Wang, Yuchen Wu 等CHI 2026 · 被引用 1 次
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal ReasoningHaozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu 等CVPR 2026 · 被引用 15 次
