SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN Training
Zhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou, Dan Feng
Abstract
Distributed deep neural networks (DNN) training systems deployed across workers have been widely used in various domains, while the communication overhead among workers for synchronizing gradient tensors often becomes the performance bottleneck. To optimize communication efficiency, state-of-the-art studies often apply both two techniques: i) gradient sparsification compression, which truncates the gradient to its largest elements to reduce the communication traffic, and ii) tensor fusion, which merges multiple gradient tensors within a fusion buffer to transmit them together to reduce the communication startup overhead. However, we find that existing studies often apply gradient sparsification after tensor fusion (we call sparsification-behind tensor fusion), which leads to a fact that a lot of fused gradient tensors are missed after the sparsification, thus impairing the convergence performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 558c2a76-9773-46cd-8133-39fe8a86ade1Related papers
- SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN TrainingZhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu et al.ICDE 2026 · 1 citation
- ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN TrainingZhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng et al.HPDC 2024 · 5 citations
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li et al.INFOCOM 2020 · 66 citations
- ZEN: Empowering Distributed Training with Sparsity-driven Data SynchronizationZhuang Wang, Zhaozhuo Xu, Jingyi Xi, Yuke Wang et al.OSDI 2025 · 3 citations
- Modeling and Optimizing the Scaling Performance in Distributed Deep Learning TrainingTing Liu, Tianhao Miao, Qinghua Wu, Zhenyu Li et al.WWW 2022 · 7 citations
