Lune

INFOCOM2024Top-tier venue

Federated Learning While Providing Model as a Service: Joint Training and Inference Optimization

Pengchao Han, Shiqiang Wang, Yang Jiao, Jianwei Huang

2024Year
19Citations
2Top-tier citations

Abstract

While providing machine learning model as a service to process users' inference requests, online applications can periodically upgrade the model utilizing newly collected data. Federated learning (FL) is beneficial for enabling the training of models across distributed clients while keeping the data locally. However, existing work has overlooked the coexistence of model training and inference under clients' limited resources. This paper focuses on the joint optimization of model training and inference to maximize inference performance at clients. Such an optimization faces several challenges. The first challenge is to characterize the clients' inference performance when clients may partially participate in FL. To resolve this challenge, we introduce a new notion of age of model (AoM) to quantify client-side model freshness, based on which we use FL's global model convergence error as an approximate measure of inference performance. The second challenge is the tight coupling among clients' decisions, including participation probability in FL, model download probability, and service rates. Toward the challenges, we propose an online problem approximation to reduce the problem complexity and optimize the resources to balance the needs of model training and inference. Experimental results demonstrate that the proposed algorithm improves the average inference accuracy by up to 12%.

Index Terms-Federated learning, model service provisioning, resource-constrained FL, model freshness, online control

• Model training resource costs in FL: Efficient FL requires active clients' participation to accelerate the training process. The participating clients in each round of FL consume communication resources to download the global model from the server and send model updates to the server. Moreover, participating clients consume computation resources for local training.

• Model inference resource costs in SP: For SP, processing users' requests requires computation resources for performing model inference. Furthermore, since clients may not participate in all rounds of FL, a client with an outdated model generally achieves worse inference performance than those who have the latest global model.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b9780f5c-5cd5-4bf6-bf3c-d5b8828f6f99

Cited by top-tier papers2

Ask how each one uses it

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines