LASER: Linear Compression in Wireless Distributed Optimization
Ashok Vardhan Makkuva, Marco Bondaschi, Thijs Vogels, Martin Jaggi, Hyeji Kim, Michael Gastpar
Abstract
Data-parallel SGD is the de facto algorithm for distributed optimization, especially for large scale machine learning. Despite its merits, communication bottleneck is one of its persistent issues. Most compression schemes to alleviate this either assume noiseless communication links, or fail to achieve good performance on practical tasks. In this paper, we close this gap and introduce LASER: LineAr CompreSsion in WirEless DistRibuted Optimization. LASER capitalizes on the inherent low-rank structure of gradients and transmits them efficiently over the noisy channels. Whilst enjoying theoretical guarantees similar to those of the classical SGD, LASER shows consistent gains over baselines on a variety of practical benchmarks. In particular, it outperforms the state-of-the-art compression schemes on challenging computer vision and GPT language modeling tasks. On the latter, we obtain 50-64% improvement in perplexity over our baselines for noisy channels. Code is available at https: //github.com/Bond1995/LASER .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- FetchSGD: Communication-Efficient Federated Learning with SketchingDaniel Rothchild, Ashwinee Panda, Enayat Ullah, Nikita Ivkin et al.ICML 2020 · 425 citations
- DRIVE: One-bit Distributed Mean EstimationShay Vargaftik, Ran Ben-Basat, Amit Portnoy, Gal Mendelson et al.NeurIPS 2021 · 82 citations
- Sparse Random Networks for Communication-Efficient Federated LearningBerivan Isik, Francesco Pase, Deniz Gündüz, Tsachy Weissman et al.ICLR 2023 · 8 citations
Related papers
- Communication-efficient Distributed Learning for Large Batch OptimizationRui Liu, Barzan MozafariICML 2022 · 9 citations
- Practical Low-Rank Communication Compression in Decentralized Deep LearningThijs Vogels, Sai Praneeth Karimireddy, Martin JaggiNeurIPS 2020 · 63 citations
- SEPARATE: A Simple Low-rank Projection for Gradient Compression in Modern Large-scale Model Training ProcessHanzhen Zhao, Xingyu Xie, Cong Fang, Zhouchen LinICLR 2025
- Linear Convergent Decentralized Optimization with CompressionXiaorui Liu, Yao Li, Rongrong Wang, Jiliang Tang et al.ICLR 2021 · 52 citations
- Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingHanyang Peng, Shuang Qin, Yue Yu, Jin Wang et al.NeurIPS 2023 · 5 citations
