Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and Inference
Huaiguang Cai, Zhi Zhou, Qianyi Huang
摘要
With edge intelligence, AI models are increasingly pushed to the edge to serve ubiquitous users. However, due to the drift of model, data, and task, AI model deployed at the edge suffers from degraded accuracy in the inference serving phase. Model retraining handles such drifts by periodically retraining the model with newly arrived data. When colocating model retraining and model inference serving for the same model on resource-limited edge servers, a fundamental challenge arises in balancing the resource allocation for model retraining and inference, aiming to maximize long-term inference accuracy. This problem is particularly difficult due to the underlying mathematical formulation being time-coupled, non-convex, and NP-hard. To address these challenges, we introduce a lightweight and explainable online approximation algorithm, named ORRIC, designed to optimize resource allocation for adaptively balancing the accuracy of model training and inference. The competitive ratio of ORRIC outperforms that of the traditional Inference-Only paradigm, especially when data drift persists for a sufficiently lengthy time. This highlights the advantages and applicable scenarios of colocating model retraining and inference. Notably, ORRIC can be translated into several heuristic algorithms for different resource environments. Experiments conducted in real scenarios validate the effectiveness of ORRIC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang 等INFOCOM 2025 · 被引用 3 次
- EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge NetworksYiming Yao, Jianwei Niu, Bin Dai, Tao RenACL 2026
它引用的顶会 Paper12
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen 等ICML 2022 · 被引用 579 次
- Continual Test-Time Domain AdaptationQin Wang, Olga Fink, Luc Van Gool, Dengxin DaiCVPR 2022 · 被引用 383 次
- Online Adaptation to Label Distribution ShiftRuihan Wu, Chuan Guo, Yi Su, Kilian Q. WeinbergerNeurIPS 2021 · 被引用 77 次
- Real-Time Video Inference on Edge Devices via Adaptive Model StreamingMehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad AlizadehICCV 2021 · 被引用 58 次
- POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and PagingShishir G. Patil, Paras Jain, Prabal Dutta, Ion Stoica 等ICML 2022 · 被引用 52 次
相关 Paper
- RESCUE: Opportunistic Online Scheduling of Model Retraining on Underutilized EdgesJianping Huang, Xiang Liu, Feng ShanINFOCOM 2026
- Online Scheduling of Edge Multiple- Model Inference with DAG Structure and RetrainingYifan Zeng, Ruiting Zhou, Lei Jiao, Renli ZhangINFOCOM 2025 · 被引用 12 次
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 被引用 34 次
- Federated Learning While Providing Model as a Service: Joint Training and Inference OptimizationPengchao Han, Shiqiang Wang, Yang Jiao, Jianwei HuangINFOCOM 2024 · 被引用 19 次
- AI in 5G: The Case of Online Distributed Transfer Learning over Edge NetworksYulan Yuan, Lei Jiao, Konglin Zhu, Xiaojun Lin 等INFOCOM 2022 · 被引用 13 次
