FLUDE: An Efficient Federated Learning Framework with Undependable Devices
Shilong Wang, Jianchun Liu, Hongli Xu, Chunming Qiao
Abstract
In a federated learning (FL) system, many devices, such as smartphones, are often undependable (e.g., frequently disconnected from WiFi) during training. Existing FL frameworks always assume a dependable environment and simply exclude undependable devices from training, leading to poor model performance and resource wastage. In this paper, we propose FLUDE to effectively deal with undependable environments. First, FLUDE assesses the dependability of devices based on the probability distribution of their historical behaviors (e.g., the likelihood of successfully completing training). Based on this assessment, FLUDE adaptively selects devices with high dependability for training. To mitigate resource wastage during the training phase, FLUDE maintains a model cache on each device, aiming to preserve the latest training state for later use in case local training on an undependable device is interrupted. Moreover, FLUDE proposes a staleness-aware strategy to judiciously distribute the global model to a subset of devices, thus significantly reducing resource wastage while maintaining model performance. We have implemented FLUDE on two physical platforms with 120 smartphones and NVIDIA Jetson devices. Extensive experiments show that FLUDE effectively improves model performance and resource efficiency of FL in undependable environments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- FedASMU: Efficient Asynchronous Federated Learning with Dynamic Staleness-Aware Model UpdateJi Liu, Juncheng Jia, Tianshi Che, Chao Huo et al.AAAI 2024 · 87 citations
- Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited StalenessHaoming Wang, Wei GaoAAAI 2025 · 3 citations
- FLuID: Mitigating Stragglers in Federated Learning using Invariant DropoutIrene Wang, Prashant J. Nair, Divya MahajanNeurIPS 2023 · 42 citations
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Aggregating Capacity in FL through Successive Layer Training for Computationally-Constrained DevicesKilian Pfeiffer, Ramin Khalili, Jörg HenkelNeurIPS 2023 · 16 citations
