LEGACY: A Lightweight Dynamic Gradient Compression Strategy for Distributed Deep Learning
Mostapha Essoullami, El Houcine Bergou, Aritra Dutta
Abstract
Distributed learning has achieved remarkable success in training deep neural networks (DNNs) on large datasets, but the communication bottleneck limits its scalability. Various compression techniques have been proposed to alleviate this limitation; however, they either use fixed parameters throughout training or rely on complex and computationally intensive methods to adapt compression parameters. Instead of the hard-to-tune hyperparameters required by adaptive compressors, this paper investigates the impact of two fundamental factors in DNN training—the layer size of the networks and their training phases—to design a simple yet efficient dynamic scheduler for any compressor, guiding the selection of compression parameters. We present a Lightweight Efficient GrAdient Compression strategyY or LEGACY, which, in theory, can work with any compression technique to produce a simple dynamic counterpart. We benchmark LEGACY on distributed and federated training, involving seven different DNN architectures, ranging from ResNet, Transformer-XL, to GPT-2, across large and challenging datasets, including ImageNet, WikiText-103, and OpenWebText. On ImageNet-1K, with an equivalent average data volume, LEGACY's dynamic compression strategies improve the Top-1 accuracy of ResNet-50 by 7-11% compared to uniform Top-0.1% compression, while on WikiText-103, the layer-based dynamic strategy reduces the perplexity of Transformer-XL by 26% relative to the same baseline. In addition, we evaluate LEGACY under constrained and federated settings, and demonstrate that it scales effectively to a 100-worker configuration while maintaining strong accuracy under aggressive compression. We publish anonymized code at: https://github.com/LEGACY-compression/LEGACY.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72959c24-cead-4ef3-808f-2fc2ac0e838eBuilds on10
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Random Reshuffling: Simple Analysis with Vast ImprovementsKonstantin Mishchenko, Ahmed Khaled, Peter RichtárikNeurIPS 2020 · 172 citations
- Rethinking gradient sparsification as total error minimizationAtal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee et al.NeurIPS 2021 · 85 citations
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui et al.NeurIPS 2020 · 81 citations
- ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive TrainingHui-Po Wang, Sebastian U. Stich, Yang He, Mario FritzICML 2022 · 70 citations
Related papers
- DC2: Delay-aware Compression Control for Distributed Machine LearningAhmed M. Abdelmoniem, Marco CaniniINFOCOM 2021 · 33 citations
- FFT-based Gradient Sparsification for the Distributed Training of Deep Neural NetworksLinnan Wang, Wei Wu, Junyu Zhang, Hang Liu et al.HPDC 2020 · 21 citations
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu et al.ICML 2023 · 7 citations
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng et al.VLDB 2022 · 39 citations
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi et al.SIGCOMM 2026 · 1 citation
