Detached Error Feedback for Distributed SGD with Random Sparsification
An Xu, Heng Huang
Abstract
The communication bottleneck has been a critical problem in large-scale distributed deep learning. In this work, we study distributed SGD with random block-wise sparsification as the gradient compressor, which is ring-allreduce compatible and highly computation-efficient but leads to inferior performance. To tackle this important issue, we improve the communication-efficient distributed SGD from a novel aspect, that is, the trade-off between the variance and second moment of the gradient. With this motivation, we propose a new detached error feedback (DEF) algorithm, which shows better convergence bound than error feedback for non-convex problems. We also propose DEF-A to accelerate the generalization of DEF at the early stages of the training, which shows better generalization bounds than DEF. Furthermore, we establish the connection between communication-efficient distributed SGD and SGD with iterate averaging (SGD-IA) for the first time. Extensive deep learning experiments show significant empirical improvement of the proposed methods under various settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcd1beb7-5821-452c-a0be-8c80f25160f6Cited by top-tier papers2
- Adversarial Weight Perturbation Improves Generalization in Graph Neural NetworksYihan Wu, Aleksandar Bojchevski, Heng HuangAAAI 2023 · 35 citations
- Ordered Momentum for Asynchronous SGDChang-Wei Shi, Yi-Rui Yang, Wu-Jun LiNeurIPS 2024 · 7 citations
Builds on18
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
- EF21: A New, Simpler, Theoretically Better, and Practically Faster Error FeedbackPeter Richtárik, Igor Sokolov, Ilyas FatkhullinNeurIPS 2021 · 219 citations
- Acceleration for Compressed Gradient Descent in Distributed and Federated OptimizationZhize Li, Dmitry Kovalev, Xun Qian, Peter RichtárikICML 2020 · 156 citations
- Rethinking gradient sparsification as total error minimizationAtal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee et al.NeurIPS 2021 · 85 citations
- On the Convergence of Communication-Efficient Local SGD for Federated LearningHongchang Gao, An Xu, Heng HuangAAAI 2021 · 66 citations
Related papers
- CSER: Communication-efficient SGD with Error ResetCong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta et al.NeurIPS 2020 · 50 citations
- Towards Faster Decentralized Stochastic Optimization with Communication CompressionRustem Islamov, Yuan Gao, Sebastian U. StichICLR 2025
- Quantized Compressive Sampling of Stochastic Gradients for Efficient Communication in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 32 citations
- Step-Ahead Error Feedback for Distributed Training with Compressed GradientAn Xu, Zhouyuan Huo, Heng HuangAAAI 2021 · 17 citations
- Analysis of Error Feedback in Federated Non-Convex Optimization with Biased Compression: Fast Convergence and Partial ParticipationXiaoyun Li, Ping LiICML 2023 · 42 citations
