Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
David Bani-Harouni, Chantal Pellegrini, Ege Özsoy, Nassir Navab, Matthias Keicher
Abstract
Clinical decision-making is a dynamic, interactive, and cyclic process where doctors have to repeatedly decide on which clinical action to perform and consider newly uncovered information for diagnosis and treatment. Large Language Models (LLMs) have the potential to support clinicians in this process, however, most applications of LLMs in clinical decision support suffer from one of two limitations: Either they assume the unrealistic scenario of immediate availability of all patient information and do not model the interactive and iterative investigation process, or they restrict themselves to the limited "out-of-the-box" capabilities of large pre-trained models without performing task-specific training. In contrast to this, we propose to model clinical decision-making for diagnosis with a hypothesis-driven uncertainty-aware language agent, LA-CDM, that converges towards a diagnosis via repeatedly requesting and interpreting relevant tests. Using a hybrid training paradigm combining supervised and reinforcement learning, we train LA-CDM with three objectives targeting critical aspects of clinical decision-making: accurate hypothesis generation, hypothesis uncertainty estimation, and efficient decision-making. We evaluate our methodology on MIMIC-CDM, a real-world dataset covering four abdominal diseases containing various clinical tests and show the benefit of explicitly training clinical decision-making for increasing diagnostic performance and efficiency. Our code is available at https://github.com/dharouni/LA-CDM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
Related papers
- Timely Clinical Diagnosis through Active Test SelectionSilas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der SchaarNeurIPS 2025 · 4 citations
- MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential DiagnosisDaniel Philip Rose, Chia-Chien Hung, Marco Lepri, Israa Alqassem et al.ACL 2025 · 14 citations
- DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent CollaborationZhihao Jia, Mingyi Jia, Junwen Duan, Jian-xin WangEMNLP 2025 · 2 citations
- Causal Discovery through Synergizing Large Language Model and Data-Driven ReasoningHuaming Du, Yujia Zheng, Baoyu Jing, Yu Zhao et al.KDD 2025 · 1 citation
- RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and TreatmentXuanzhong Chen, Ye Jin, Xiaohao Mao, Lun Wang et al.AAAI 2026 · 10 citations
