Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference
Huaiguang Cai, Zhi Zhou, Qianyi Huang
Abstract
With edge intelligence, AI models are increasingly pushed to the edge to serve ubiquitous users. However, due to the drift of model, data, and task, AI model deployed at the edge suffers from degraded accuracy in the inference serving phase. Model retraining handles such drifts by periodically retraining the model with newly arrived data. When colocating model retraining and model inference serving for the same model on resource-limited edge servers, a fundamental challenge arises in balancing the resource allocation for model retraining and inference, aiming to maximize long-term inference accuracy. This problem is particularly difficult due to the underlying mathematical formulation being time-coupled, non-convex, and NP-hard. To address these challenges, we introduce a lightweight and explainable online approximation algorithm, named ORRIC, designed to optimize resource allocation for adaptively balancing the accuracy of model training and inference. The competitive ratio of ORRIC outperforms that of the traditional Inference-Only paradigm, especially when data drift persists for a sufficiently lengthy time. This highlights the advantages and applicable scenarios of colocating model retraining and inference. Notably, ORRIC can be translated into several heuristic algorithms for different resource environments. Experiments conducted in real scenarios validate the effectiveness of ORRIC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bcacdcf-ed89-41b2-8574-e32c73741cfeCited by top-tier papers2
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang et al.INFOCOM 2025 · 3 citations
- EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge NetworksYiming Yao, Jianwei Niu, Bin Dai, Tao RenACL 2026
Builds on12
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
- Continual Test-Time Domain AdaptationQin Wang, Olga Fink, Luc Van Gool, Dengxin DaiCVPR 2022 · 383 citations
- Online Adaptation to Label Distribution ShiftRuihan Wu, Chuan Guo, Yi Su, Kilian Q. WeinbergerNeurIPS 2021 · 77 citations
- Real-Time Video Inference on Edge Devices via Adaptive Model StreamingMehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad AlizadehICCV 2021 · 58 citations
- POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and PagingShishir G. Patil, Paras Jain, Prabal Dutta, Ion Stoica et al.ICML 2022 · 52 citations
Related papers
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
- Online Scheduling of Edge Multiple- Model Inference with DAG Structure and RetrainingYifan Zeng, Ruiting Zhou, Lei Jiao, Renli ZhangINFOCOM 2025 · 12 citations
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 34 citations
- Federated Learning While Providing Model as a Service: Joint Training and Inference OptimizationPengchao Han, Shiqiang Wang, Yang Jiao, Jianwei HuangINFOCOM 2024 · 19 citations
- AI in 5G: The Case of Online Distributed Transfer Learning over Edge NetworksYulan Yuan, Lei Jiao, Konglin Zhu, Xiaojun Lin et al.INFOCOM 2022 · 13 citations
