Lune

INFOCOM2024顶会

Federated Learning While Providing Model as a Service: Joint Training and Inference Optimization

Pengchao Han, Shiqiang Wang, Yang Jiao, Jianwei Huang

2024年份
19被引次数
2顶会引用

摘要

While providing machine learning model as a service to process users' inference requests, online applications can periodically upgrade the model utilizing newly collected data. Federated learning (FL) is beneficial for enabling the training of models across distributed clients while keeping the data locally. However, existing work has overlooked the coexistence of model training and inference under clients' limited resources. This paper focuses on the joint optimization of model training and inference to maximize inference performance at clients. Such an optimization faces several challenges. The first challenge is to characterize the clients' inference performance when clients may partially participate in FL. To resolve this challenge, we introduce a new notion of age of model (AoM) to quantify client-side model freshness, based on which we use FL's global model convergence error as an approximate measure of inference performance. The second challenge is the tight coupling among clients' decisions, including participation probability in FL, model download probability, and service rates. Toward the challenges, we propose an online problem approximation to reduce the problem complexity and optimize the resources to balance the needs of model training and inference. Experimental results demonstrate that the proposed algorithm improves the average inference accuracy by up to 12%.

Index Terms-Federated learning, model service provisioning, resource-constrained FL, model freshness, online control

• Model training resource costs in FL: Efficient FL requires active clients' participation to accelerate the training process. The participating clients in each round of FL consume communication resources to download the global model from the server and send model updates to the server. Moreover, participating clients consume computation resources for local training.

• Model inference resource costs in SP: For SP, processing users' requests requires computation resources for performing model inference. Furthermore, since clients may not participate in all rounds of FL, a client with an outdated model generally achieves worse inference performance than those who have the latest global model.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext b9780f5c-5cd5-4bf6-bf3c-d5b8828f6f99

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖