AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge Servers
Sudipta Saha Shubha, Haiying Shen
摘要
Various audio and video applications rely on multiple deep neural network (DNN) models deployed on edge servers to conduct inference with ms-level latency service-level-objectives (SLOs). To avoid accuracy decreases caused by data drift, continual retraining is necessary. However, this poses a challenge for GPU resource allocation to satisfy the tight SLOs while maintaining high accuracy in this scenario. There has been no research devoted to tackling this issue. In this paper, we conducted trace-based experimental analysis in this particular scenario, which shows that different models have varying degrees of impact from data drift, incremental retraining (proposed by us that retrains certain samples before inference) and early-exit model structures can help increase accuracy, and the interdependencies among tasks may lead to significant CPU-GPU memory communications. Leveraging these unique observations, we propose a data drift Adaptive scheduler for accurate and SLO-guaranteed Inference serving at edge servers (AdaInf). AdaInf uses incremental retraining and allocates GPU amount among applications based on their SLOs. For each application, it splits GPU time between retraining and inference to satisfy its SLO, and then allocates GPU time among retraining tasks based on their impact degrees. In addition, AdaInf proposes strategies that leverage the job features in this scenario to reduce the impact of CPU-GPU memory communications on latency. Our real trace-driven experimental evaluation shows that AdaInf can increase accuracy by up to 21% and reduce SLO violations by up to 54% compared to existing methods. Achieving similar accuracy as AdaInf requires 4× more GPU resources on the edge server for the existing method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFSZiyang Zhang, Yang Zhao, Ming-Ching Chang, Changyao Lin 等AAAI 2025 · 被引用 4 次
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 被引用 4 次
- On-Demand Container Partitioning for Distributed MLGiovanni Bartolomeo, Navidreza Asadi, Wolfgang Kellerer, Jörg Ott 等USENIX ATC 2025 · 被引用 3 次
- Characterizing Mobile SoC for Accelerating Heterogeneous LLM InferenceLe Chen, Dahu Feng, Erhu Feng, Yingrui Wang 等SOSP 2025 · 被引用 3 次
相关 Paper
- Online Scheduling of Edge Multiple- Model Inference with DAG Structure and RetrainingYifan Zeng, Ruiting Zhou, Lei Jiao, Renli ZhangINFOCOM 2025 · 被引用 12 次
- Ekya: Continuous Learning of Video Analytics Models on Edge Compute ServersRomil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang 等NSDI 2022
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon 等EuroSys 2026 · 被引用 1 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
