Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AI
Zhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang, Fangming Liu, Xu Chen
摘要
Edge inference leverages edge computing devices for the last mile delivery of artificial intelligence (AI) services. To meet the latency requirements while overcoming the resource limitations, edge inference systems deploy lightweight - typically compressed - DNN models. However, due to data drift during deployment, these compressed edge models often fail to deliver satisfactory and stable inference accuracy. To address this issue, we propose a novel edge inference serving paradigm called Federated Inference. This approach, based on ensemble learning, groups multiple edge workers to form an ensemble, enhancing inference accuracy. A key challenge in Federated Inference is maximizing ensemble accuracy while adhering to resource budgets and Service Level Objectives (SLOs). The dynamic nature of the environment and the NP-hardness of the optimization problem add to the complexity. To address these challenges, we propose an online model ensemble framework that integrates online learning with approximate optimization, offering a theoretically rigorous and computationally efficient solution. We have implemented a prototype of our framework and, through extensive test-bed evaluations, demonstrate that it improves average inference accuracy by.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 被引用 329 次
- Resource-Efficient Federated Learning with Hierarchical Aggregation in Edge ComputingZhiyuan Wang, Hongli Xu, Jianchun Liu, He Huang 等INFOCOM 2021 · 被引用 216 次
- Billion-scale federated learning on mobile clients: a submodel design with tunable privacyChaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua 等MobiCom 2020 · 被引用 114 次
- Distributed Inference with Deep Learning Models across Heterogeneous Edge DevicesChenghao Hu, Baochun LiINFOCOM 2022 · 被引用 81 次
- Heterogeneity-Aware Federated Learning with Adaptive Client Selection and Gradient CompressionZhida Jiang, Yang Xu, Hongli Xu, Zhiyuan Wang 等INFOCOM 2023 · 被引用 43 次
相关 Paper
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
- Jellyfish: Timely Inference Serving for Dynamic Edge NetworksVinod Nigade, Pablo Bauszat, Henri E. Bal, Lin WangRTSS 2022 · 被引用 40 次
- AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge ServersSudipta Saha Shubha, Haiying ShenSIGCOMM 2023 · 被引用 34 次
- Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy DistillationYaofo Chen, Shuaicheng Niu, Yaowei Wang, Shoukai Xu 等ICLR 2024 · 被引用 18 次
- Federated Learning While Providing Model as a Service: Joint Training and Inference OptimizationPengchao Han, Shiqiang Wang, Yang Jiao, Jianwei HuangINFOCOM 2024 · 被引用 19 次
