Understanding the effects of data parallelism and sparsity on neural network training
Namhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr, Martin Jaggi
Abstract
We study two factors in neural network training: data parallelism and sparsity; here, data parallelism means processing training data in parallel using distributed systems (or equivalently increasing batch size), so that training can be accelerated; for sparsity, we refer to pruning parameters in a neural network model, so as to reduce computational and memory cost. Despite their promising benefits, however, understanding of their effects on neural network training remains elusive. In this work, we first measure these effects rigorously by conducting extensive experiments while tuning all metaparameters involved in the optimization. As a result, we find across various workloads of data set, network model, and optimization algorithm that there exists a general scaling trend between batch size and number of training steps to convergence for the effect of data parallelism, and further, difficulty of training under sparsity. Then, we develop a theoretical analysis based on the convergence properties of stochastic gradient methods and smoothness of the optimization landscape, which illustrates the observed phenomena precisely and generally, establishing a better account of the effects of data parallelism and sparsity on neural network training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Meta-Learning Sparse Implicit Neural RepresentationsJaeho Lee, Jihoon Tack, Namhoon Lee, Jinwoo ShinNeurIPS 2021 · 60 citations
- SAFE: Finding Sparse and Flat Minima to Improve PruningDongyeop Lee, Kwanhee Lee, Jinseok Chung, Namhoon LeeICML 2025
Builds on4
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 174 citations
Related papers
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 25 citations
- Convergence-Aware Neural Network TrainingHyungjun Oh, Yongseung Yu, Giha Ryu, Gunjoo Ahn et al.DAC 2020 · 4 citations
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho et al.AAAI 2020
- Pruning's Effect on Generalization Through the Lens of Training and RegularizationTian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle et al.NeurIPS 2022 · 40 citations
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi et al.CVPR 2020
