Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
Taeyoon Kwon, Kai Tzu-iunn Ong, Dongjin Kang, Seungjun Moon, Jeong Ryong Lee, Dosik Hwang, Beomseok Sohn, Yongsik Sim, Dongha Lee, Jinyoung Yeo
Abstract
Machine reasoning has made great progress in recent years owing to large language models (LLMs). In the clinical domain, however, most NLP-driven projects mainly focus on clinical classification or reading comprehension, and under-explore clinical reasoning for disease diagnosis due to the expensive rationale annotation with clinicians. In this work, we present a "reasoning-aware" diagnosis framework that rationalizes the diagnostic process via prompt-based learning in a time- and labor-efficient manner, and learns to reason over the prompt-generated rationales. Specifically, we address the clinical reasoning for disease diagnosis, where the LLM generates diagnostic rationales providing its insight on presented patient data and the reasoning path towards the diagnosis, namely Clinical Chain-of-Thought (Clinical CoT). We empirically demonstrate LLMs/LMs' ability of clinical reasoning via extensive experiments and analyses on both rationale generation and disease diagnosis in various settings. We further propose a novel set of criteria for evaluating machine-generated rationales' potential for real-world clinical settings, facilitating and benefiting future research in this area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7c140a7-c8a4-4e1c-9ce8-f1698b46cf51Cited by top-tier papers7
- Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language ModelsShuai Niu, Jing Ma, Hongzhan Lin, Liang Bai et al.ACL 2025 · 5 citations
- SINCon: Mitigate LLM-Generated Malicious Message Injection Attack for Rumor DetectionMingqing Zhang, Qiang Liu, Xiang Tao, Shu Wu et al.ACL 2025 · 2 citations
- Knowledgeable Language Models as Black-Box Optimizers for Personalized MedicineMichael S. Yao, Osbert Bastani, Alma Andersson, Tommaso Biancalani et al.ICLR 2026
- BoxLM: Unifying Structures and Semantics of Medical Concepts for Diagnosis Prediction in HealthcareYanchao Tan, Hang Lv, Yunfei Zhan, Guofang Ma et al.ICML 2025
- Enhancing Graph Of Thought: Enhancing Prompts with LLM Rationales and Dynamic Temperature ControlSunguk Shin, Youngjoon KimICLR 2025
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
Related papers
- M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image UnderstandingJuntao Jiang, Jiangning Zhang, Yali Bi, Jinsheng Bai et al.ICLR 2026 · 3 citations
- OncoCoT: A Temporal-causal Chain-of-Thought Dataset for Oncologic Decision-MakingPeiru Yang, Yudong Li, Shiting Wang, Xinyi Liu et al.AAAI 2026
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative SearchHaoran Sun, Yankai Jiang, Wenjie Lou, Yujie Zhang et al.NeurIPS 2025 · 16 citations
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought BenchmarkGui Wang, YongSong Zhou, Kaijun Deng, Wooi Ping Cheah et al.CVPR 2026
