DynamiQ: Accelerating Gradient Synchronization using Compressed Multi-hop All-reduce
Wenchen Han, Shay Vargaftik, Michael Mitzenmacher, Ran Ben Basat
摘要
Multi-hop all-reduce is the de facto backbone of large model training. As the training scale increases, the network often becomes a bottleneck, motivating the reduction of the volume of transmitted data. Accordingly, recent systems have demonstrated significant acceleration of the training process using gradient quantization. However, these systems are not optimized for multi-hop aggregation, where entries are partially summed multiple times along their aggregation topology.
We present DynamiQ, a quantization framework that bridges the gap between quantization best practices and multi-hop aggregation. DynamiQ introduces novel techniques to better represent partial sums, co-designed with a decompress-accumulate-recompress fused kernel to facilitate fast execution.
We extend PyTorch DDP to support DynamiQ over NCCL P2P, and across different LLMs, tasks, and scales, we demonstrate consistent improvement of up to 34.2% over the best among state-of-theart methods such as Omni-Reduce, THC, and emerging standards such as MXFP4, MXFP6, and MXFP8. Further, DynamiQ is the only evaluated method that consistently reaches near-baseline accuracy (e.g., 99.9% of the BF16 baseline) and does so while significantly accelerating the training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
相关 Paper
- AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenXiang Lin, HuangJunTao, LuHan Zhang, Lilaiyi 等ICML 2026
- Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AIMikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman 等SC 2024 · 被引用 15 次
- DeMo: Decoupled Momentum OptimizationBowen Peng, Lizhang Chen, Baiyu Su, Jeffrey Quesnelle 等ICLR 2026 · 被引用 7 次
- Better Together: Jointly Optimizing ML Collective Scheduling and Execution Planning using SYNDICATEKshiteej Mahajan, Ching-Hsiang Chu, Srinivas Sridharan, Aditya AkellaNSDI 2023 · 被引用 43 次
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 被引用 1 次
