Online Scheduling of Edge Multiple- Model Inference with DAG Structure and Retraining
Yifan Zeng, Ruiting Zhou, Lei Jiao, Renli Zhang
Abstract
Edge inference applications are becoming increasingly complex and composed of multiple models. The dependency of models is modeled by a Directed Acyclic Graph (DAG). The accuracy of the edge model is easily affected by data drift. Retraining is employed to sustain the inference accuracy of models. But the introduction of retraining complicates the inter-task dependency of the inference request. Moreover, model retraining prolongs the inference completion time. The accuracy improvement and the latency increment under different retraining configurations necessitate a trade-off between inference accu-racy and request completion time. In this paper, we investigate multiple-model inference with retraining, aiming to maximize inference accuracy while minimizing request completion time. Through experimental analysis, we observed that retraining can enhance model inference accuracy in a short time. We represent the retraining tasks of models as nodes in the DAG of the inference request and then construct a unified DAG structure for both retraining and inference tasks. We first propose a Single Request Scheduling Algorithm (SRS) with a theoretical performance guarantee to select the optimal retraining configuration for each model under edge resource constraints and jointly schedule retraining and inference tasks. Subsequently, we extend SRS to a Multiple Requests Scheduling Algorithm (MRS) to address scheduling in a more general online multi-request scenario. The experiments on an edge system indicate that compared to existing methods, MRS can enhance the inference accuracy by 25 % while reducing the request completion time by 45 %.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ca3577aa-c1f9-4007-9e90-e3b3c26a6a29Cited by top-tier papers2
- A First Look at Operational RAN Updates and Their Impact on Carrier Traffic Demands and PredictionAntonio Boiano, Nadezda Chukhno, Zbigniew Smoreda, Alessandro Enrico Cesare Redondi et al.INFOCOM 2026 · 1 citation
- PARD: Enhancing Goodput for Inference Pipeline via Proactive Request DroppingZhixin Zhao, Yitao Hu, Simin Chen, Mingfang Ji et al.EuroSys 2026
Related papers
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 34 citations
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 10 citations
- Ekya: Continuous Learning of Video Analytics Models on Edge Compute ServersRomil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang et al.NSDI 2022
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon et al.EuroSys 2026 · 1 citation
