HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units
Shenda Hong, Yanbo Xu, Alind Khare, Satria Priambada, Kevin O. Maher, Alaa Aljiffry, Jimeng Sun, Alexey Tumanov
Abstract
Deep learning models have achieved expert-level performance in healthcare with an exclusive focus on training accurate models. However, in many clinical environments such as intensive care unit (ICU), real-time model serving is equally if not more important than accuracy, because in ICU patient care is simultaneously more urgent and more expensive. Clinical decisions and their timeliness, therefore, directly affect both the patient outcome and the cost of care. To make timely decisions, we argue the underlying serving system must be latency-aware. To compound the challenge, health analytic applications often require a combination of models instead of a single model, to better specialize individual models for different targets, multi-modal data, different prediction windows, and potentially personalized predictions. To address these challenges, we propose HOLMES---an online model ensemble serving framework for healthcare applications. HOLMES dynamically identifies the best performing set of models to ensemble for highest accuracy, while also satisfying sub-second latency constraints on end-to-end prediction. We demonstrate that HOLMES is able to navigate the accuracy/latency tradeoff efficiently, compose the ensemble, and serve the model ensemble pipeline, scaling to simultaneously streaming data from 100 patients, each producing waveform data at 250 Hz. HOLMES outperforms the conventional offline batch-processed inference for the same clinical task in terms of accuracy and latency (by order of magnitude). HOLMES is tested on risk prediction task on pediatric cardio ICU data with above 95% prediction accuracy and sub-second latency on 64-bed simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd4e5a21-2c4a-442c-b3e5-de2ae9a2d5acCited by top-tier papers7
- Efficient Architecture Search for Diverse TasksJunhong Shen, Mikhail Khodak, Ameet TalwalkarNeurIPS 2022 · 42 citations
- TusoAI: Agentic Optimization for Scientific MethodsAlistair Turcan, Kexin Huang, Lei Li, Martin J. ZhangICLR 2026 · 3 citations
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame ProjectionsBerken Utku Demirel, Christian HolzNeurIPS 2025 · 1 citation
- Dist Loss: Enhancing Regression in Few-Shot Region through Distribution Distance ConstraintGuangkun Nie, Gongzheng Tang, Shenda HongICLR 2025
- Shifting the Paradigm: A Diffeomorphism Between Time Series Data Manifolds for Achieving Shift-Invariancy in Deep LearningBerken Utku Demirel, Christian HolzICLR 2025
Related papers
- Cocktail: A Multidimensional Optimization for Model Serving in CloudJashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma et al.NSDI 2022
- UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage ClassificationYanbo Xu, Alind Khare, Glenn Matlin, Monish Ramadoss et al.NeurIPS 2022 · 4 citations
- Efficient Deep Ensemble Inference via Query Difficulty-dependent Task SchedulingZichong Li, Lan Zhang, Mu Yuan, Miaohui Song et al.ICDE 2023 · 4 citations
- DeepAlerts: Deep Learning Based Multi-Horizon Alerts for Clinical Deterioration on Oncology Hospital WardsDingwen Li, Patrick G. Lyons, Chenyang Lu, Marin KollefAAAI 2020 · 18 citations
- Joint Model and Data Adaptation for Cloud Inference ServingJingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He et al.RTSS 2021 · 19 citations
