ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
Liwen Sun, Abhineet Agarwal, Aaron Kornblith, Bin Yu, Chenyan Xiong
Abstract
In the emergency department (ED), patients undergo triage and multiple laboratory tests before diagnosis. This time-consuming process causes ED crowding which impacts patient mortality, medical errors, staff burnout, etc. This work proposes (time) cost-effective diagnostic assistance that leverages artificial intelligence systems to help ED clinicians make efficient and accurate diagnoses. In collaboration with ED clinicians, we use public patient data to curate MIMIC-ED-Assist, a benchmark for AI systems to suggest laboratory tests that minimize wait time while accurately predicting critical outcomes such as death. With MIMIC-ED-Assist, we develop ED-Copilot which sequentially suggests patient-specific laboratory tests and makes diagnostic predictions. ED-Copilot employs a pre-trained bio-medical language model to encode patient information and uses reinforcement learning to minimize ED wait time and maximize prediction accuracy. On MIMIC-ED-Assist, ED-Copilot improves prediction accuracy over baselines while halving average wait time from four hours to two hours. ED-Copilot can also effectively personalize treatment recommendations based on patient severity, further highlighting its potential as a diagnostic assistant. Since MIMIC-ED-Assist is a retrospective benchmark, ED-Copilot is restricted to recommend only observed tests. We show ED-Copilot achieves competitive performance without this restriction as the maximum allowed time increases. Our code is available at https: //github.com/cxcscmu/ED-Copilot .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e0a554d-fa70-4b8c-8641-342df8aaa434Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Hierarchical Shrinkage: Improving the accuracy and interpretability of tree-based modelsAbhineet Agarwal, Yan Shuo Tan, Omer Ronen, Chandan Singh et al.ICML 2022 · 37 citations
- Language Models are Weak LearnersHariharan Manikandan, Yiding Jiang, J. Zico KolterNeurIPS 2023 · 32 citations
- Deep Reinforcement Learning for Cost-Effective Medical DiagnosisZheng Yu, Yikuan Li, Joseph C. Kim, Kaixuan Huang et al.ICLR 2023 · 13 citations
Related papers
- Timely Clinical Diagnosis through Active Test SelectionSilas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der SchaarNeurIPS 2025 · 4 citations
- The Impact of Auxiliary Patient Data on Automated Chest X-Ray Report Generation and How to Incorporate ItAaron Nicolson, Shengyao Zhuang, Jason Dowling, Bevan KoopmanACL 2025 · 6 citations
- With, not For: Co-Designing a Patient-Facing AI Companion Concept for the Emergency Department Waiting AreaJacobe Klein, Peter Sörries, Yasemin Mutlugil, Ceenu George et al.CHI 2026 · 4 citations
- EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context ReasoningHengyu Zhang, Xuyun Zhang, Pengxiang Zhan, Linhao Luo et al.KDD 2026
- Salus: Strategic Diagnostic Testing for Complex Diagnosis via Multi-Agent Reinforcement LearningShuohao Gao, Xuanzhong Chen, Lingxiao Luo, Zilin Ding et al.ICML 2026
