Sign bit is enough: a learning synchronization framework for multi-hop all-reduce with ultimate compression
Feijie Wu, Shiqi He, Song Guo, Zhihao Qu, Haozhao Wang, Weihua Zhuang, Jie Zhang
摘要
Traditional one-bit compressed stochastic gradient descent can not be directly employed in multi-hop all-reduce, a widely adopted distributed training paradigm in network-intensive high-performance computing systems such as public clouds. According to our theoretical findings, due to the cascading compression, the training process has considerable deterioration on the convergence performance. To overcome this limitation, we implement a sign-bit compression-based learning synchronization framework, Marsit. It prevents cascading compression via an elaborate bit-wise operation for unbiased sign aggregation and its specific global compensation mechanism for mitigating compression deviation. The proposed framework retains the same theoretical convergence rate as non-compression mechanisms. Experimental results demonstrate that Marsit reduces up to 35% training time while preserving the same accuracy as training without compression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel ExtractionFeijie Wu, Xingchen Wang, Yaqing Wang, Tianci Liu 等NeurIPS 2024 · 被引用 47 次
- Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingHanyang Peng, Shuang Qin, Yue Yu, Jin Wang 等NeurIPS 2023 · 被引用 5 次
- FedFetch: Faster Federated Learning with Adaptive Downstream PrefetchingQifan Yan, Andrew Liu, Shiqi He, Mathias Lécuyer 等INFOCOM 2025 · 被引用 2 次
- Towards Privacy-Preserving and Heterogeneity-aware Split Federated Learning via Probabilistic MaskingXingchen Wang, Feijie Wu, Chenglin Miao, Tianchun Li 等KDD 2026
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence SpeedHanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari 等ICML 2021 · 被引用 106 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- Stochastic Sign Descent Methods: New Algorithms and Better TheoryMher Safaryan, Peter RichtárikICML 2021 · 被引用 70 次
相关 Paper
- DUO: No Compromise to Accuracy DegradationJinda Jia, Cong Xie, Hanlin Lu, Fanjiang Ye 等NeurIPS 2025
- IntSGD: Adaptive Floatless Compression of Stochastic GradientsKonstantin Mishchenko, Bokun Wang, Dmitry Kovalev, Peter RichtárikICLR 2022 · 被引用 19 次
- On Distributed Adaptive Optimization with Gradient CompressionXiaoyun Li, Belhal Karimi, Ping LiICLR 2022 · 被引用 34 次
- On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep LearningAritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 等AAAI 2020
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable DevicesMax Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, Gennady PekhimenkoNeurIPS 2021 · 被引用 59 次
