LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI Hospitals
Zonghai Yao, Hong Yu
Abstract
This survey reviews LLM-based multi-agent systems for clinical and healthcare workflows, including diagnosis, triage, consultation, discharge, mental health, and EHR-linked decision support. We define AI hospitals as workflow-level clinical systems in which agents take explicit roles, hand off shared state, use EHR-or guideline-grounded tools, and operate with safety gates and audit-ready logs. We argue that these systems should be compared at the workflow level, rather than only by model components or end-task accuracy, because clinical action, evidence, and accountability are expressed through state transitions and handoffs. We organize the literature through a workflow-level taxonomy covering roles and handoffs, memory and evidence, tools, and reasoning, control, and escalation. We further synthe-size major workflow settings and task families, introduce a four-layer evaluation stack spanning safety, process, outcome, and operations, and connect model capabilities to workflow observables relevant to deployment. Finally, we present Integration Readiness Levels (IRL1-IRL6), task-level instrumentation requirements, and recurring workflow failure modes as a practical framework for comparing, evaluating, and deploying clinical LLM agents and AI hospitals. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowZiyue Wang, Junde Wu, Linghan Cai, Chang Han Low et al.ICLR 2026 · 84 citations
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett et al.ICML 2024 · 82 citations
- Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen et al.ACL 2025 · 64 citations
- DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert KnowledgeBufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu et al.UbiComp 2025 · 60 citations
Related papers
- Learning to Be a Doctor: Searching for Effective Medical Agent ArchitecturesYangyang Zhuang, Wenjia Jiang, Jiayu Zhang, Ze Yang et al.ACM MM 2025 · 1 citation
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin et al.CHI 2026 · 2 citations
- LiveClin: A Live Clinical Benchmark without LeakageXidong Wang, Guo shuqi, Yue Shen, Junying Chen et al.ICLR 2026 · 4 citations
- CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI WorkflowsAmit Giloni, Chiara Picardi, Roy Betser, Shamik Bose et al.EMNLP 2025
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz et al.EMNLP 2025 · 2 citations
