Federated Learning While Providing Model as a Service: Joint Training and Inference Optimization
Pengchao Han, Shiqiang Wang, Yang Jiao, Jianwei Huang
Abstract
While providing machine learning model as a service to process users' inference requests, online applications can periodically upgrade the model utilizing newly collected data. Federated learning (FL) is beneficial for enabling the training of models across distributed clients while keeping the data locally. However, existing work has overlooked the coexistence of model training and inference under clients' limited resources. This paper focuses on the joint optimization of model training and inference to maximize inference performance at clients. Such an optimization faces several challenges. The first challenge is to characterize the clients' inference performance when clients may partially participate in FL. To resolve this challenge, we introduce a new notion of age of model (AoM) to quantify client-side model freshness, based on which we use FL's global model convergence error as an approximate measure of inference performance. The second challenge is the tight coupling among clients' decisions, including participation probability in FL, model download probability, and service rates. Toward the challenges, we propose an online problem approximation to reduce the problem complexity and optimize the resources to balance the needs of model training and inference. Experimental results demonstrate that the proposed algorithm improves the average inference accuracy by up to 12%.
Index Terms-Federated learning, model service provisioning, resource-constrained FL, model freshness, online control
• Model training resource costs in FL: Efficient FL requires active clients' participation to accelerate the training process. The participating clients in each round of FL consume communication resources to download the global model from the server and send model updates to the server. Moreover, participating clients consume computation resources for local training.
• Model inference resource costs in SP: For SP, processing users' requests requires computation resources for performing model inference. Furthermore, since clients may not participate in all rounds of FL, a client with an outdated model generally achieves worse inference performance than those who have the latest global model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9780f5c-5cd5-4bf6-bf3c-d5b8828f6f99Cited by top-tier papers2
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- OPTION: An Online Pricing Strategy for Asynchronous Federated Learning Against Free-Riding AttacksBangqi Pan, Jianfeng Lu, Shuqin Cao, Xiao Zhang et al.AAAI 2026
Builds on16
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated LearningHaibo Yang, Minghong Fang, Jia LiuICLR 2021 · 310 citations
- Clustered Sampling: Low-Variance and Improved Representativity for Clients Selection in Federated LearningYann Fraboni, Richard Vidal, Laetitia Kameni, Marco LorenziICML 2021 · 249 citations
- MARINA: Faster Non-Convex Distributed Learning with CompressionEduard Gorbunov, Konstantin Burlachenko, Zhize Li, Peter RichtárikICML 2021 · 129 citations
Related papers
- SplitGP: Achieving Both Generalization and Personalization in Federated LearningDong-Jun Han, Do-Yeon Kim, Minseok Choi, Christopher G. Brinton et al.INFOCOM 2023 · 43 citations
- Federated Learning with Flexible ControlShiqiang Wang, Jake B. Perazzone, Mingyue Ji, Kevin S. ChanINFOCOM 2023 · 30 citations
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 10 citations
- To Store or Not? Online Data Selection for Federated Learning with Limited StorageChen Gong, Zhenzhe Zheng, Fan Wu, Yunfeng Shao et al.WWW 2023 · 28 citations
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang et al.INFOCOM 2025 · 3 citations
