Efficient Deep Ensemble Inference via Query Difficulty-dependent Task Scheduling
Zichong Li, Lan Zhang, Mu Yuan, Miaohui Song, Qi Song
摘要
Deep ensemble learning has been widely adopted to boost accuracy through combing outputs from multiple deep models prepared for the same task. However, the extra computation and memory cost it entails could impose an unacceptably high deadline miss rate in latency-sensitive tasks. Conventional approaches, including ensemble selection, focus on accuracy while ignoring deadline constraints, and thus cannot smartly cope with bursty query traffic and queries with different hardness. This paper explores redundancy in deep ensemble model inference and presents Schemble, a query difficulty-dependent task scheduling framework. Schemble treats ensemble inference progress as multiple base model inference tasks and schedules tasks for queries based on their difficulty and queuing status. We evaluate Schemble on real-world datasets, considering intelligent Q&A system, video analysis and image retrieval as the running applications. Experimental results show that Schemble achieves a 5× lower deadline miss rate and improves the accuracy by 30.8% given deadline constraints.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Model Selection for Latency-Critical Inference ServingDaniel Mendoza, Francisco Romero, Caroline TrippelEuroSys 2024 · 被引用 16 次
- MLink: Linking Black-Box Models for Collaborative Multi-Model InferenceMu Yuan, Lan Zhang, Xiang-Yang LiAAAI 2022 · 被引用 10 次
- Zygarde: Time-Sensitive On-Device Deep Inference and Adaptation on Intermittently-Powered SystemsBashima Islam, Shahriar NirjonUbiComp 2020 · 被引用 68 次
- Real-Time Multitasking of Deep Neural Networks With Nvidia TensorrtFederico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. ButtazzoRTSS 2025 · 被引用 1 次
- HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care UnitsShenda Hong, Yanbo Xu, Alind Khare, Satria Priambada 等KDD 2020 · 被引用 81 次
