PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
Yisu Wang, Ruilong Wu, Xinjiao Li, Dirk Kutscher
摘要
Large-scale deep neural networks (DNN) exhibit excellent performance for various tasks. As DNNs and datasets grow, distributed training becomes extremely time-consuming and demands larger clusters. A main bottleneck is the resulting gradient aggregation overhead. While gradient compression and sparse collective communication techniques are commonly employed to alleviate network load, many gradient compression schemes do not achieve acceleration of the training process while also preserving accuracy. This paper introduces PacTrain, a novel framework that accelerates distributed training by combining pruning with sparse gradient compression. Active pruning of the neural network makes the model weights and gradients sparse. By ensuring the global knowledge of the gradient sparsity among all distributed training workers, we can perform lightweight compression communication without harming accuracy. We show that the PacTrain compression scheme achieves a near-optimal compression strategy while remaining compatible with the allreduce primitive. Experimental evaluations show that PacTrain improves training throughput by 1.25 to compared to state-of-the-art compression-enabled systems for representative vision and language models training tasks under bandwidth-constrained conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao 等NeurIPS 2023 · 被引用 293 次
相关 Paper
- THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic CompressionMinghao Li, Ran Ben Basat, Shay Vargaftik, ChonLam Lao 等NSDI 2024 · 被引用 44 次
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui 等NeurIPS 2020 · 被引用 81 次
- Communication Efficient SGD via Gradient Sampling With Bayes PriorLiuyihan Song, Kang Zhao, Pan Pan, Yu Liu 等CVPR 2021
- Gradient Compression Supercharged High-Performance Data Parallel DNN TrainingYouhui Bai, Cheng Li, Quan Zhou, Jun Yi 等SOSP 2021 · 被引用 36 次
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 被引用 57 次
