Towards a Better Theoretical Understanding of Independent Subnetwork Training
Egor Shulgin, Peter Richtárik
Abstract
Modern advancements in large-scale machine learning would be impossible without the paradigm of data-parallel distributed computing. Since distributed computing with large-scale models imparts excessive pressure on communication channels, significant recent research has been directed toward co-designing communication compression strategies and training algorithms with the goal of reducing communication costs. While pure data parallelism allows better data scaling, it suffers from poor model scaling properties. Indeed, compute nodes are severely limited by memory constraints, preventing further increases in model size. For this reason, the latest achievements in training giant neural network models also rely on some form of model parallelism. In this work, we take a closer theoretical look at Independent Subnetwork Training (IST), which is a recently proposed and highly effective technique for solving the aforementioned problems. We identify fundamental differences between IST and alternative approaches, such as distributed methods with compressed communication, and provide a precise analysis of its optimization performance on a quadratic model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca3578c0-620a-4aff-8662-c72d6ef866aeCited by top-tier papers4
- FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel ExtractionFeijie Wu, Xingchen Wang, Yaqing Wang, Tianci Liu et al.NeurIPS 2024 · 47 citations
- FedP3: Federated Personalized and Privacy-friendly Network Pruning under Model HeterogeneityKai Yi, Nidham Gazagnadou, Peter Richtárik, Lingjuan LyuICLR 2024 · 18 citations
- Sparser, Better, Deeper, Stronger: Improving Static Sparse Training with Exact Orthogonal InitializationAleksandra Nowak, Lukasz Gniecki, Filip Szatkowski, Jacek TaborICML 2024 · 4 citations
- MAST: model-agnostic sparsified trainingYury Demidovich, Grigory Malinovsky, Egor Shulgin, Peter RichtárikICLR 2025
Builds on12
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutSamuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis et al.NeurIPS 2021 · 390 citations
- FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model ExtractionSamiul Alam, Luyang Liu, Ming Yan, Mi ZhangNeurIPS 2022 · 261 citations
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 200 citations
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous ClientsEnmao Diao, Jie Ding, Vahid TarokhICLR 2021 · 179 citations
Related papers
- Distributed Learning of Fully Connected Neural Networks using Independent Subnet TrainingBinhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang et al.VLDB 2022 · 42 citations
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 5 citations
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho et al.AAAI 2020
- Understanding the effects of data parallelism and sparsity on neural network trainingNamhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr, Martin JaggiICLR 2021 · 8 citations
- Communication-efficient Distributed Learning for Large Batch OptimizationRui Liu, Barzan MozafariICML 2022 · 9 citations
