Online Scheduling of Edge Multiple- Model Inference with DAG Structure and Retraining
Yifan Zeng, Ruiting Zhou, Lei Jiao, Renli Zhang
摘要
Edge inference applications are becoming increasingly complex and composed of multiple models. The dependency of models is modeled by a Directed Acyclic Graph (DAG). The accuracy of the edge model is easily affected by data drift. Retraining is employed to sustain the inference accuracy of models. But the introduction of retraining complicates the inter-task dependency of the inference request. Moreover, model retraining prolongs the inference completion time. The accuracy improvement and the latency increment under different retraining configurations necessitate a trade-off between inference accu-racy and request completion time. In this paper, we investigate multiple-model inference with retraining, aiming to maximize inference accuracy while minimizing request completion time. Through experimental analysis, we observed that retraining can enhance model inference accuracy in a short time. We represent the retraining tasks of models as nodes in the DAG of the inference request and then construct a unified DAG structure for both retraining and inference tasks. We first propose a Single Request Scheduling Algorithm (SRS) with a theoretical performance guarantee to select the optimal retraining configuration for each model under edge resource constraints and jointly schedule retraining and inference tasks. Subsequently, we extend SRS to a Multiple Requests Scheduling Algorithm (MRS) to address scheduling in a more general online multi-request scenario. The experiments on an edge system indicate that compared to existing methods, MRS can enhance the inference accuracy by 25 % while reducing the request completion time by 45 %.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- A First Look at Operational RAN Updates and Their Impact on Carrier Traffic Demands and PredictionAntonio Boiano, Nadezda Chukhno, Zbigniew Smoreda, Alessandro Enrico Cesare Redondi 等INFOCOM 2026 · 被引用 1 次
- PARD: Enhancing Goodput for Inference Pipeline via Proactive Request DroppingZhixin Zhao, Yitao Hu, Simin Chen, Mingfang Ji 等EuroSys 2026
相关 Paper
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 被引用 34 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
- Ekya: Continuous Learning of Video Analytics Models on Edge Compute ServersRomil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang 等NSDI 2022
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon 等EuroSys 2026 · 被引用 1 次
