Moniqua: Modulo Quantized Communication in Decentralized SGD
Yucheng Lu, Christopher De Sa
Abstract
Running Stochastic Gradient Descent (SGD) in a decentralized fashion has shown promising results. In this paper we propose Moniqua, a technique that allows decentralized SGD to use quantized communication. We prove in theory that Moniqua communicates a provably bounded number of bits per iteration, while converging at the same asymptotic rate as the original algorithm does with full-precision communication. Moniqua improves upon prior works in that it (1) requires zero additional memory, (2) works with 1-bit quantization, and ( 3 ) is applicable to a variety of decentralized algorithms. We demonstrate empirically that Moniqua converges faster with respect to wall clock time than other quantized decentralized algorithms. We also show that Moniqua is robust to very low bit-budgets, allowing 1-bit-per-parameter communication without compromising validation accuracy when training ResNet20 and ResNet110 on CIFAR10.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3620cbb-27c2-4f0a-b5a9-1692a49fd7f4Cited by top-tier papers14
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 118 citations
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 95 citations
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li et al.NeurIPS 2021 · 61 citations
- Cross-Gradient Aggregation for Decentralized Learning from Non-IID DataYasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu et al.ICML 2021 · 61 citations
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen et al.ICML 2023 · 60 citations
Related papers
- Resource-Aware Decentralized Learning with Rate-Adaptive QuantizationJing Qiao, Yu Liu, Yuan Yuan, Yifei Zou et al.INFOCOM 2026
- Quantized Decentralized Stochastic Learning over Directed GraphsHossein Taheri, Aryan Mokhtari, Hamed Hassani, Ramtin PedarsaniICML 2020 · 59 citations
- SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign DecodingChanho Park, Namyoon LeeICML 2024 · 5 citations
- Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingHanyang Peng, Shuang Qin, Yue Yu, Jin Wang et al.NeurIPS 2023 · 5 citations
- DSGD-CECA: Decentralized SGD with Communication-Optimal Exact Consensus AlgorithmLisang Ding, Kexin Jin, Bicheng Ying, Kun Yuan et al.ICML 2023 · 12 citations
