Deep Reinforcement Learning for Cost-Effective Medical Diagnosis
Zheng Yu, Yikuan Li, Joseph C. Kim, Kaixuan Huang, Yuan Luo, Mengdi Wang
Abstract
Dynamic diagnosis is desirable when medical tests are costly or time-consuming. In this work, we use reinforcement learning (RL) to find a dynamic policy that selects lab test panels sequentially based on previous observations, ensuring accurate testing at a low cost. Clinical diagnostic data are often highly imbalanced; therefore, we aim to maximize the score instead of the error rate. However, optimizing the non-concave score is not a classic RL problem, thus invalidates standard RL methods. To remedy this issue, we develop a reward shaping approach, leveraging properties of the score and duality of policy optimization, to provably find the set of all Pareto-optimal policies for budget-constrained score maximization. To handle the combinatorially complex state space, we propose a Semi-Model-based Deep Diagnosis Policy Optimization (SM-DDPO) framework that is compatible with end-to-end training and online learning. SM-DDPO is tested on diverse clinical tasks: ferritin abnormality detection, sepsis mortality prediction, and acute kidney injury diagnosis. Experiments with real-world data validate that SM-DDPO trains efficiently and identifies all Pareto-front solutions. Across all tasks, SM-DDPO is able to achieve state-of-the-art diagnosis accuracy (in some cases higher than conventional methods) with up to reduction in testing cost. The code is available at [https://github.com/Zheng321/Deep-Reinforcement-Learning-for-Cost-Effective-Medical-Diagnosis].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08764ebe-a11d-43c8-9e0d-33b9d4d19637Cited by top-tier papers4
- Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement LearningDavid Bani-Harouni, Chantal Pellegrini, Ege Özsoy, Nassir Navab et al.ICLR 2026 · 16 citations
- ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic AssistanceLiwen Sun, Abhineet Agarwal, Aaron Kornblith, Bin Yu et al.ICML 2024 · 7 citations
- Limited Query Graph Connectivity TestMingyu Guo, Jialiang Li, Aneta Neumann, Frank Neumann et al.AAAI 2024 · 7 citations
- Timely Clinical Diagnosis through Active Test SelectionSilas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der SchaarNeurIPS 2025 · 4 citations
Builds on2
- Active Feature Acquisition with Generative Surrogate ModelsYang Li, Junier OlivaICML 2021 · 52 citations
- Reinforcement Learning with State Observation Costs in Action-Contingent Noiselessly Observable Markov Decision ProcessesHyunji Alex Nam, Scott L. Fleming, Emma BrunskillNeurIPS 2021 · 34 citations
Related papers
- Acquisition Conditioned Oracle for Nongreedy Active Feature AcquisitionMichael Valancius, Max Lennon, Junier OlivaICML 2024 · 7 citations
- Salus: Strategic Diagnostic Testing for Complex Diagnosis via Multi-Agent Reinforcement LearningShuohao Gao, Xuanzhong Chen, Lingxiao Luo, Zilin Ding et al.ICML 2026
- Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱Lu Wang, Wenchao Yu, Xiaofeng He, Wei Cheng et al.WWW 2020 · 33 citations
- Dialogue Based Disease Screening Through Domain Customized Reinforcement LearningZhuo Liu, Yanxuan Li, Xingzhi Sun, Fei Wang et al.KDD 2021 · 8 citations
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingDavid Dai, Peilin Chen, Chanakya Ekbote, Paul Pu LiangNeurIPS 2025 · 48 citations
