ErrorCompensatedX: error compensation for variance reduced algorithms
Hanlin Tang, Yao Li, Ji Liu, Ming Yan
Abstract
Communication cost is one major bottleneck for the scalability for distributed learning. One approach to reduce the communication cost is to compress the gradient during communication. However, directly compressing the gradient decelerates the convergence speed, and the resulting algorithm may diverge for biased compression. Recent work addressed this problem for stochastic gradient descent by adding back the compression error from the previous step. This idea was further extended to one class of variance reduced algorithms, where the variance of the stochastic gradient is reduced by taking a moving average over all history gradients. However, our analysis shows that just adding the previous step's compression error, as done in existing work, does not fully compensate the compression error. So, we propose ErrorCompensatedX, which uses the compression error from the previous two steps. We show that ErrorCompensatedX can achieve the same asymptotic convergence rate with the training without compression. Moreover, we provide a unified theoretical analysis framework for this class of variance reduced algorithms, with or without error compensation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Detached Error Feedback for Distributed SGD with Random SparsificationAn Xu, Heng HuangICML 2022 · 12 citations
- Error Feedback under (L0, L1)-Smoothness: Normalization and MomentumSarit Khirirat, Abdurakhmon Sadiev, Artem Riabinin, Eduard Gorbunov et al.NeurIPS 2025 · 10 citations
- A Tight Theory of Error Feedback Algorithms in Distributed OptimizationDaniel Berg Thomsen, Adrien Taylor, Aymeric DieuleveutICML 2026
Builds on6
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow MomentumJianyu Wang, Vinayak Tantia, Nicolas Ballas, Michael G. RabbatICLR 2020 · 220 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- Linearly Converging Error Compensated SGDEduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, Peter RichtárikNeurIPS 2020 · 90 citations
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui et al.NeurIPS 2020 · 81 citations
- CSER: Communication-efficient SGD with Error ResetCong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta et al.NeurIPS 2020 · 50 citations
Related papers
- On Distributed Adaptive Optimization with Gradient CompressionXiaoyun Li, Belhal Karimi, Ping LiICLR 2022 · 34 citations
- Quantized Compressive Sampling of Stochastic Gradients for Efficient Communication in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 32 citations
- Distributed Bilevel Optimization with Communication CompressionYutong He, Jie Hu, Xinmeng Huang, Songtao Lu et al.ICML 2024 · 2 citations
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho et al.AAAI 2020
- Error Compensated Distributed SGD Can Be AcceleratedXun Qian, Peter Richtárik, Tong ZhangNeurIPS 2021 · 65 citations
