LLMs (Almost) Never Abstain Under Medical Uncertainty
Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Gianluca Moro
Abstract
Medical multiple-choice question answering (MCQA) benchmarks implicitly assume that large language models (LLMs) should always commit to an answer. However, in clinical practice, uncertainty is pervasive and abstaining is often the safest action. We introduce MedQAbstain, a benchmark explicitly designed to evaluate medical abstention under uncertainty. MedQAbstain repurposes standard medical MCQA datasets by removing the gold answer and introducing an explicit "I abstain" option, framed as a safety-critical decision with clinical consequences. The benchmark supports systematic analysis across abstention regimes, distractor complexity, and input modalities, and elicits self-reported model confidence to study calibration. Across all settings, we find that state-of-the-art LLMs systematically overcommit, rarely abstaining even when the question itself is hidden. These results reveal a fundamental mismatch between LLM behavior and clinical norms, highlighting abstention as a critical but overlooked dimension of medical decision-making evaluation. 1 * Equal contribution (co-first authors).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 865b261c-ba8a-4e8d-8219-711d888075ddBuilds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Reasoning Models Better Express Their ConfidenceDongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim et al.NeurIPS 2025 · 77 citations
- Discriminative Marginalized Probabilistic Neural Method for Multi-Document Summarization of Medical LiteratureGianluca Moro, Luca Ragazzi, Lorenzo Valgimigli, Davide FreddiACL 2022 · 42 citations
Related papers
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model UncertaintyJingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang et al.ACL 2026 · 1 citation
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical ReasoningXilin Dang, Kexin Chen, Xiaorui Su, Ayush Noori et al.ICLR 2026 · 6 citations
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsShrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang et al.EMNLP 2025 · 8 citations
- Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language ModelsVy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He et al.AAAI 2026 · 2 citations
