Distributed Learning of Fully Connected Neural Networks using Independent Subnet Training
Binhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang, Anastasios Kyrillidis, Chris Jermaine
Abstract
Distributed machine learning (ML) can bring more computational resources to bear than single-machine learning, thus enabling reductions in training time. Distributed learning partitions models and data over many machines, allowing model and dataset sizes beyond the available compute power and memory of a single machine. In practice though, distributed ML is challenging when distribution is mandatory, rather than chosen by the practitioner. In such scenarios, data could unavoidably be separated among workers due to limited memory capacity per worker or even because of data privacy issues. There, existing distributed methods will utterly fail due to dominant transfer costs across workers, or do not even apply. We propose a new approach to distributed fully connected neural network learning, called independent subnet training (IST), to handle these cases. In IST, the original network is decomposed into a set of narrow subnetworks with the same depth. These subnetworks are then trained locally before parameters are exchanged to produce new subnets and the training cycle repeats. Such a naturally "model parallel" approach limits memory usage by storing only a portion of network parameters on each device. Additionally, no requirements exist for sharing data between workers (i.e., subnet training is local and independent) and communication volume and frequency are reduced by decomposing the original network into independent subnets. These properties of IST can cope with issues due to distributed data, slow interconnects, or limited device memory, making IST a suitable approach for cases of mandatory distribution. We show experimentally that IST results in training times that are much lower than common distributed learning approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d92b0ee1-520c-4879-8eb9-2e77df3e78a2Cited by top-tier papers8
- Helios: Heterogeneity-Aware Federated Learning with Dynamically Balanced CollaborationZirui Xu, Fuxun Yu, Jinjun Xiong, Xiang ChenDAC 2021 · 50 citations
- RSC: Accelerate Graph Neural Networks Training via Randomized Sparse ComputationsZirui Liu, Shengyuan Chen, Kaixiong Zhou, Daochen Zha et al.ICML 2023 · 24 citations
- FedP3: Federated Personalized and Privacy-friendly Network Pruning under Model HeterogeneityKai Yi, Nidham Gazagnadou, Peter Richtárik, Lingjuan LyuICLR 2024 · 18 citations
- Federated Learning Over Images: Vertical Decompositions and Pre-Trained Backbones Are Difficult to BeatErdong Hu, Yuxin Tang, Anastasios Kyrillidis, Chris JermaineICCV 2023 · 13 citations
- Towards a Better Theoretical Understanding of Independent Subnetwork TrainingEgor Shulgin, Peter RichtárikICML 2024 · 8 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- TiFL: A Tier-based Federated Learning SystemZheng Chai, Ahsan Ali, Syed Zawad, Stacey Truex et al.HPDC 2020 · 330 citations
- Inefficiency of K-FAC for Large Batch Size TrainingLinjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao et al.AAAI 2020 · 24 citations
Related papers
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Aggregating Capacity in FL through Successive Layer Training for Computationally-Constrained DevicesKilian Pfeiffer, Ramin Khalili, Jörg HenkelNeurIPS 2023 · 16 citations
- Device-Wise Federated Network PruningShangqian Gao, Junyi Li, Zeyu Zhang, Yanfu Zhang et al.CVPR 2024
- Workflow Optimization for Parallel Split LearningJoana Tirana, Dimitra Tsigkari, George Iosifidis, Dimitris ChatzopoulosINFOCOM 2024 · 14 citations
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous ClientsEnmao Diao, Jie Ding, Vahid TarokhICLR 2021 · 179 citations
