FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks
Linnan Wang, Wei Wu, Junyu Zhang, Hang Liu, George Bosilca, Maurice Herlihy, Rodrigo Fonseca
Abstract
The performance and efficiency of distributed training of Deep Neural Networks (DNN) highly depend on the performance of gradient averaging among participating processes, a step bound by communication costs. There are two major approaches to reduce communication overhead: overlap communications with computations (lossless), or reduce communications (lossy). The lossless solution works well for linear neural architectures, e.g. VGG, AlexNet, but more recent networks such as ResNet and Inception limit the opportunity for such overlapping. Therefore, approaches that reduce the amount of data (lossy) become more suitable. In this paper, we present a novel, explainable lossy method that sparsifies gradients in the frequency domain, in addition to a new range-based float point representation to quantize and further compress gradients. These dynamic techniques strike a balance between compression ratio, accuracy, and computational overhead, and are optimized to maximize performance in heterogeneous environments.
Unlike existing works that strive for a higher compression ratio, we stress the robustness of our methods, and provide guidance to recover accuracy from failures. To achieve this, we prove how the FFT sparsification affects the convergence and accuracy, and show that our method is guaranteed to converge using a diminishing 𝜃 in training. Reducing 𝜃 can also be used to recover accuracy from the failure. Compared to STOA lossy methods, e.g., QSGD, TernGrad, and Top-k sparsification, our approach incurs less approximation error, thereby better in both the wall-time and accuracy. On an 8 GPUs, InfiniBand interconnected cluster, our techniques effectively accelerate AlexNet training up to 2.26x to the baseline of no compression, and 1.31x to QSGD, 1.25x to Terngrad and 1.47x to Top-K sparsification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 271a088d-9213-45ce-9c5a-06f3f4aef38aCited by top-tier papers3
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 57 citations
- Dr. Top-k: delegate-centric Top-k on GPUsAnil Gaihre, Da Zheng, Scott Weitze, Lingda Li et al.SC 2021 · 16 citations
- Distributed Parallel Gradient Stacking(DPGS): Solving Whole Slide Image Stacking Challenge in Multi-Instance LearningBoyuan Wu, Zefeng Wang, Xianwei Lin, Jiachun Xu et al.ICML 2025
Related papers
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He et al.ICDE 2023 · 9 citations
- COMPSO: Optimizing Gradient Compression for Distributed Training with Second-Order OptimizersBaixi Sun, Weijin Liu, J. Gregory Pauloski, Jiannan Tian et al.PPoPP 2025 · 8 citations
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho et al.AAAI 2020
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui et al.NeurIPS 2020 · 81 citations
- Communication Efficient SGD via Gradient Sampling With Bayes PriorLiuyihan Song, Kang Zhao, Pan Pan, Yu Liu et al.CVPR 2021
