SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN Training
Zhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu, Wenxiang Zhou, Xinjue Zheng, Dan Feng
Abstract
Distributed deep neural networks (DNN) training systems deployed across multiple workers have been widely used to accelerate the training of large models and datasets, while the communication overhead for synchronizing data (i.e., gradient tensor) among workers often becomes a performance bottleneck. To improve communication efficiency, two key techniques are often adopted: i) tensor fusion, which merges multiple tensors to transmit them together to reduce the communication startup overhead, and ii) gradient sparsification compression, which transmits only the largest gradient elements to reduce communication traffic. Recent studies focus on combining tensor fusion and sparsification: a straightforward way is to perform sparsification on each tensor before fusion (per-tensor sparsification), but it incurs significant sparsification overhead; alternatively, state-ofthe-art works first merge multiple tensors (i.e., fusion) and then perform sparsification on them to reduce sparsification overhead (multi-tensor sparsification). However, we observe that multi-tensor sparsification causes many tensors to be missing, leading to a significant decrease in convergence accuracy. In this paper, we revisit per-tensor sparsification for no tensor missing and propose a new selective sparsification mechanism tailored for fused tensors. This mechanism strategically selects a subset of fused tensors: the selected ones undergo individual sparsification, while the others remain unsparsified. This design achieves both low sparsification overhead (via selective sparsification) and high convergence performance (without tensor missing), which we term SSFusion. On top of SSFusion, we propose an efficient sparsification offloading scheme that offloads GPU-based gradient sparsification to the CPU to further speed up sparsification, and an interleaved communication scheme that improves communication efficiency through fusion separation. Evaluations on cloud and local clusters show that SSFusion improves the training throughput by over stateof-the-art solutions, while maintaining approximately the same convergence accuracy as the non-sparsification baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 71043c12-6494-41b2-af2d-9cfdf4604566Related papers
- SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingZhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou et al.HPDC 2025 · 1 citation
- ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN TrainingZhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng et al.HPDC 2024 · 5 citations
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li et al.INFOCOM 2020 · 66 citations
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin et al.INFOCOM 2023 · 4 citations
- ZEN: Empowering Distributed Training with Sparsity-driven Data SynchronizationZhuang Wang, Zhaozhuo Xu, Jingyi Xi, Yuke Wang et al.OSDI 2025 · 3 citations
