Moniqua: Modulo Quantized Communication in Decentralized SGD
Yucheng Lu, Christopher De Sa
摘要
Running Stochastic Gradient Descent (SGD) in a decentralized fashion has shown promising results. In this paper we propose Moniqua, a technique that allows decentralized SGD to use quantized communication. We prove in theory that Moniqua communicates a provably bounded number of bits per iteration, while converging at the same asymptotic rate as the original algorithm does with full-precision communication. Moniqua improves upon prior works in that it (1) requires zero additional memory, (2) works with 1-bit quantization, and ( 3 ) is applicable to a variety of decentralized algorithms. We demonstrate empirically that Moniqua converges faster with respect to wall clock time than other quantized decentralized algorithms. We also show that Moniqua is robust to very low bit-budgets, allowing 1-bit-per-parameter communication without compromising validation accuracy when training ResNet20 and ResNet110 on CIFAR10.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li 等NeurIPS 2021 · 被引用 61 次
- Cross-Gradient Aggregation for Decentralized Learning from Non-IID DataYasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu 等ICML 2021 · 被引用 61 次
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen 等ICML 2023 · 被引用 60 次
相关 Paper
- Resource-Aware Decentralized Learning with Rate-Adaptive QuantizationJing Qiao, Yu Liu, Yuan Yuan, Yifei Zou 等INFOCOM 2026
- Quantized Decentralized Stochastic Learning over Directed GraphsHossein Taheri, Aryan Mokhtari, Hamed Hassani, Ramtin PedarsaniICML 2020 · 被引用 59 次
- SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign DecodingChanho Park, Namyoon LeeICML 2024 · 被引用 5 次
- Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingHanyang Peng, Shuang Qin, Yue Yu, Jin Wang 等NeurIPS 2023 · 被引用 5 次
- DSGD-CECA: Decentralized SGD with Communication-Optimal Exact Consensus AlgorithmLisang Ding, Kexin Jin, Bicheng Ying, Kun Yuan 等ICML 2023 · 被引用 12 次
