Enhancing Federated Learning with In-Cloud Unlabeled Data
Lun Wang, Yang Xu, Hongli Xu, Jianchun Liu, Zhiyuan Wang, Liusheng Huang
Abstract
Federated learning (FL) has been widely applied to collaboratively train deep learning (DL) models on massive end devices (i.e., clients). Due to the limited storage capacity and high labeling cost, there are always insufficient data stored and annotated on each client. Conversely, in cloud datacenters, there exist large-scale unlabeled data, which are easy to collect from public access (e.g., social media). Herein, upon the federated semi-supervised learning (FSSL) technology, we propose the Ada-FedSemi system, which leverages both on-device labeled data and in-cloud unlabeled data to boost the performance of DL models. Given the limited communication and massive quantity of the clients, in each training round, we decide to select partial clients to participate in FL, and their local models are aggregated by the parameter server (PS) to produce pseudo-labels for the unlabeled data, which are utilized to enhance the global model. Considering that the number of participating clients and the quality of pseudo-labels will have a significant impact on the training performance (e.g., efficiency and accuracy), we introduce a multi-armed bandit (MAB) based online algorithm to adaptively determine the participating fraction and confidence threshold during federated model training. Extensive experiments on benchmark models and datasets show that, given the same resource budget, the model trained by Ada-FedSemi achieves 3%-14.8 % higher test accuracy than that of the baseline methods. Besides, when achieving the same test accuracy, Ada-FedSemi saves up to 48% training cost, compared with the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48c6064e-8559-42d8-bae3-b4eff6e8268dCited by top-tier papers1
Ask how each one uses itBuilds on14
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Optimizing Federated Learning on Non-IID Data with Reinforcement LearningHao Wang, Zakhary Kaplan, Di Niu, Baochun LiINFOCOM 2020 · 1,002 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
Related papers
- SemiFL: Semi-Supervised Federated Learning for Unlabeled Clients with Alternate TrainingEnmao Diao, Jie Ding, Vahid TarokhNeurIPS 2022 · 130 citations
- SemiDFL: A Semi-Supervised Paradigm for Decentralized Federated LearningXinyang Liu, Pengchao Han, Xuan Li, Bo LiuAAAI 2025 · 3 citations
- Class Balanced Adaptive Pseudo Labeling for Federated Semi-Supervised LearningMing Li, Qingli Li, Yan WangCVPR 2023
- (FL)2: Overcoming Few Labels in Federated Semi-Supervised LearningSeungjoo Lee, Thanh-Long V. Le, Jaemin Shin, Sung-Ju LeeNeurIPS 2024 · 15 citations
- Local or Global: Selective Knowledge Assimilation for Federated Learning with Limited LabelsYae Jee Cho, Gauri Joshi, Dimitrios DimitriadisICCV 2023 · 11 citations
