DISTREAL: Distributed Resource-Aware Learning in Heterogeneous Systems
Martin Rapp, Ramin Khalili, Kilian Pfeiffer, Jörg Henkel
Abstract
We study the problem of distributed training of neural networks (NNs) on devices with heterogeneous, limited, and time-varying availability of computational resources. We present an adaptive, resource-aware, on-device learning mechanism, DISTREAL, which is able to fully and efficiently utilize the available resources on devices in a distributed manner, increasing the convergence speed. This is achieved with a dropout mechanism that dynamically adjusts the computational complexity of training an NN by randomly dropping filters of convolutional layers of the model. Our main contribution is the introduction of a design space exploration (DSE) technique, which finds Pareto-optimal per-layer dropout vectors with respect to resource requirements and convergence speed of the training. Applying this technique, each device is able to dynamically select the dropout vector that fits its available resource without requiring any assistance from the server. We implement our solution in a federated learning (FL) system, where the availability of computational resources varies both between devices and over time, and show through extensive evaluation that we are able to significantly increase the convergence speed over the state of the art without compromising on the final accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- FLuID: Mitigating Stragglers in Federated Learning using Invariant DropoutIrene Wang, Prashant J. Nair, Divya MahajanNeurIPS 2023 · 42 citations
- Aggregating Capacity in FL through Successive Layer Training for Computationally-Constrained DevicesKilian Pfeiffer, Ramin Khalili, Jörg HenkelNeurIPS 2023 · 16 citations
- LATTE: Layer Algorithm-aware Training Time Estimation for Heterogeneous Federated LearningKun Wang, Zimu Zhou, Zhenjiang LiMobiCom 2024 · 13 citations
- EchoPFL: Asynchronous Personalized Federated Learning on Mobile Devices with On-Demand Staleness ControlXiaochen Li, Sicong Liu, Zimu Zhou, Bin Guo et al.UbiComp 2024 · 11 citations
- Moss: Proxy Model-based Full-Weight Aggregation in Federated Learning with Heterogeneous ModelsYifeng Cai, Ziqi Zhang, Ding Li, Yao Guo et al.UbiComp 2025
Builds on3
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutSamuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis et al.NeurIPS 2021 · 390 citations
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous ClientsEnmao Diao, Jie Ding, Vahid TarokhICLR 2021 · 179 citations
Related papers
- WHALE-FL: Wireless and Heterogeneity Aware Latency Efficient Federated Learning over Mobile Devices via Adaptive Subnetwork SchedulingHuai-An Su, Jiaxiang Geng, Liang Li, Xiaoqi Qin et al.AAAI 2025 · 1 citation
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Workflow Optimization for Parallel Split LearningJoana Tirana, Dimitra Tsigkari, George Iosifidis, Dimitris ChatzopoulosINFOCOM 2024 · 14 citations
- Decentralized Sporadic Federated Learning: A Unified Algorithmic Framework with Convergence GuaranteesShahryar Zehtabi, Dong-Jun Han, Rohit Parasnis, Seyyedali Hosseinalipour et al.ICLR 2025
- Federated Learning Based on Dynamic RegularizationDurmus Alp Emre Acar, Yue Zhao, Ramon Matas Navarro, Matthew Mattina et al.ICLR 2021 · 114 citations
