Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed Optimization
Mher Safaryan, Filip Hanzely, Peter Richtárik
Abstract
Large scale distributed optimization has become the default tool for the training of supervised machine learning models with a large number of parameters and training data. Recent advancements in the field provide several mechanisms for speeding up the training, including compressed communication, variance reduction and acceleration. However, none of these methods is capable of exploiting the inherently rich data-dependent smoothness structure of the local losses beyond standard smoothness constants. In this paper, we argue that when training supervised models, smoothness matricesinformation-rich generalizations of the ubiquitous smoothness constants-can and should be exploited for further dramatic gains, both in theory and practice. In order to further alleviate the communication burden inherent in distributed optimization, we propose a novel communication sparsification strategy that can take full advantage of the smoothness matrices associated with local losses. To showcase the power of this tool, we describe how our sparsification technique can be adapted to three distributed optimization algorithms-DCGD [Khirirat et al., 2018] , DIANA [Mishchenko et al., 2019] and ADIANA [Li et al., 2020] -yielding significant savings in terms of communication complexity. The new methods always outperform the baselines, often dramatically so.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Fine-tuning Language Models over Slow Networks using Activation Quantization with GuaranteesJue Wang, Binhang Yuan, Luka Rimanic, Yongjun He et al.NeurIPS 2022 · 37 citations
- FedP3: Federated Personalized and Privacy-friendly Network Pruning under Model HeterogeneityKai Yi, Nidham Gazagnadou, Peter Richtárik, Lingjuan LyuICLR 2024 · 18 citations
- The Power of Extrapolation in Federated LearningHanmin Li, Kirill Acharya, Peter RichtárikNeurIPS 2024 · 16 citations
- Searching for Optimal Per-Coordinate Step-sizes with Multidimensional BacktrackingFrederik Kunstner, Victor Sanches Portella, Mark Schmidt, Nicholas J. A. HarveyNeurIPS 2023 · 14 citations
- Knowledge Distillation Performs Partial Variance ReductionMher Safaryan, Alexandra Peste, Dan AlistarhNeurIPS 2023 · 14 citations
Builds on6
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Is Local SGD Better than Minibatch SGD?Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai et al.ICML 2020 · 277 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Acceleration for Compressed Gradient Descent in Distributed and Federated OptimizationZhize Li, Dmitry Kovalev, Xun Qian, Peter RichtárikICML 2020 · 156 citations
- Linearly Converging Error Compensated SGDEduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, Peter RichtárikNeurIPS 2020 · 90 citations
Related papers
- Theoretically Better and Numerically Faster Distributed Optimization with Smoothness-Aware Quantization TechniquesBokun Wang, Mher Safaryan, Peter RichtárikNeurIPS 2022 · 13 citations
- Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?Yutong He, Xinmeng Huang, Kun YuanNeurIPS 2023 · 25 citations
- CANITA: Faster Rates for Distributed Convex Optimization with Communication CompressionZhize Li, Peter RichtárikNeurIPS 2021 · 36 citations
- Rethinking gradient sparsification as total error minimizationAtal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee et al.NeurIPS 2021 · 85 citations
- Communication-efficient Distributed Learning for Large Batch OptimizationRui Liu, Barzan MozafariICML 2022 · 9 citations
