Communication Efficient SGD via Gradient Sampling With Bayes Prior
Liuyihan Song, Kang Zhao, Pan Pan, Yu Liu, Yingya Zhang, Yinghui Xu, Rong Jin
摘要
Gradient compression has been widely adopted in dataparallel distributed training of deep neural networks to reduce communication overhead. Some literatures have demonstrated that large gradients are more important than small ones because they contain more information, such as Top-k compressor. Other mainstream methods, like random-k compressor and gradient quantization, usually treat all gradients equally. Different from all of them, we regard large and small gradients selection as the exploitation and exploration of gradient information, respectively. And we find taking both of them into consideration is the key to boost the final accuracy. So, we propose a novel gradient compressor: Gradient Sampling with Bayes Prior in this paper. Specifically, we sample important/large gradients based on the global gradient distribution, which is periodically updated across multiple workers. Then we introduce Bayes Prior into distribution model to further explore the gradients. We prove the convergence of our method for smooth non-convex problems in the distributed system. Compared with methods that running after high compression ratio at the expense of accuracy, we pursue no loss of accuracy and the actual acceleration benefit in practice. Experimental comparisons on a variety of computer vision tasks (e.g. image classification and object detection) and backbones (ResNet, MobileNetV2, InceptionV3 and AlexNet) show that our approach outperforms the stateof-the-art techniques in terms of both speed and accuracy, with the limitation of 100× compression ratio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Detached Error Feedback for Distributed SGD with Random SparsificationAn Xu, Heng HuangICML 2022 · 被引用 12 次
- JointSQ: Joint Sparsification-Quantization for Distributed LearningWeiying Xie, Haowei Li, Jitao Ma, Yunsong Li 等CVPR 2024 · 被引用 9 次
它引用的顶会 Paper8
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- Adaptive Gradient Quantization for Data-Parallel SGDFartash Faghri, Iman Tabrizian, Ilia Markov, Dan Alistarh 等NeurIPS 2020 · 被引用 108 次
- CSER: Communication-efficient SGD with Error ResetCong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta 等NeurIPS 2020 · 被引用 50 次
- Variance Reduction With Sparse GradientsMelih Elibol, Lihua Lei, Michael I. JordanICLR 2020 · 被引用 25 次
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 等AAAI 2020
相关 Paper
- SK-Gradient: Efficient Communication for Distributed Machine Learning with Data SketchJie Gui, Yuchen Song, Zezhou Wang, Chenhong He 等ICDE 2023 · 被引用 9 次
- SwitchTop-k: Scaling Top-k Compression on Programmable SwitchesYijun Li, Jiawei Huang, Jingling Liu, Zhaoyi Li 等KDD 2025
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui 等NeurIPS 2020 · 被引用 81 次
- DAGC: Data-Aware Adaptive Gradient CompressionRongwei Lu, Jiajun Song, Bin Chen, Laizhong Cui 等INFOCOM 2023 · 被引用 12 次
- Error Compensated Distributed SGD Can Be AcceleratedXun Qian, Peter Richtárik, Tong ZhangNeurIPS 2021 · 被引用 65 次
