DC2: Delay-aware Compression Control for Distributed Machine Learning
Ahmed M. Abdelmoniem, Marco Canini
摘要
Distributed training performs data-parallel training of DNN models which is a necessity for increasingly complex models and large datasets. Recent works are identifying major communication bottlenecks in distributed training. These works seek possible opportunities to speed-up the training in systems supporting distributed ML workloads. As communication reduction, compression techniques are proposed to speed up this communication phase. However, compression comes at the cost of reduced model accuracy, especially when compression is applied arbitrarily. Instead, we advocate a more controlled use of compression and propose DC2, a delay-aware compression control mechanism. DC2 couples compression control and network delays in applying compression adaptively. DC2 not only compensates for network variations but can also strike a better trade-off between training speed and accuracy. DC2 is implemented as a drop-in module to the communication library used by the ML toolkit and can operate in a variety of network settings. We empirically evaluate DC2 in network environments exhibiting low and high delay variations. Our evaluation of different popular CNN models and datasets shows that DC2 improves training speed-ups of up to 41× and 5.3 × over baselines with no-compression and uniform compression, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Optimal Rate Adaption in Federated Learning with Compressed CommunicationsLaizhong Cui, Xiaoxin Su, Yipeng Zhou, Jiangchuan LiuINFOCOM 2022 · 被引用 61 次
- Federated Learning with Flexible ControlShiqiang Wang, Jake B. Perazzone, Mingyue Ji, Kevin S. ChanINFOCOM 2023 · 被引用 30 次
- Accelerating Model Training in Multi-cluster Environments with Consumer-grade GPUsHwijoon Lim, Juncheol Ye, Sangeetha Abdu Jyothi, Dongsu HanSIGCOMM 2024 · 被引用 22 次
- DAGC: Data-Aware Adaptive Gradient CompressionRongwei Lu, Jiajun Song, Bin Chen, Laizhong Cui 等INFOCOM 2023 · 被引用 12 次
- Expediting In-Network Federated Learning by Voting-Based Consensus Model CompressionXiaoxin Su, Yipeng Zhou, Laizhong Cui, Song GuoINFOCOM 2024 · 被引用 7 次
它引用的顶会 Paper2
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 等AAAI 2020
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
相关 Paper
- LEGACY: A Lightweight Dynamic Gradient Compression Strategy for Distributed Deep LearningMostapha Essoullami, El Houcine Bergou, Aritra DuttaICLR 2026
- SAPipe: Staleness-Aware Pipeline for Data Parallel DNN TrainingYangrui Chen, Cong Xie, Meng Ma, Juncheng Gu 等NeurIPS 2022 · 被引用 24 次
- Gradient Compression Supercharged High-Performance Data Parallel DNN TrainingYouhui Bai, Cheng Li, Quan Zhou, Jun Yi 等SOSP 2021 · 被引用 36 次
- A Flexible Framework for Communication-Efficient Machine LearningSarit Khirirat, Sindri Magnússon, Arda Aytekin, Mikael JohanssonAAAI 2021 · 被引用 16 次
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep LearningYisu Wang, Ruilong Wu, Xinjiao Li, Dirk KutscherDAC 2025 · 被引用 3 次
