Lune

ICDE2026顶会

SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN Training

Zhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu, Wenxiang Zhou, Xinjue Zheng, Dan Feng

2026年份
1被引次数

摘要

Distributed deep neural networks (DNN) training systems deployed across multiple workers have been widely used to accelerate the training of large models and datasets, while the communication overhead for synchronizing data (i.e., gradient tensor) among workers often becomes a performance bottleneck. To improve communication efficiency, two key techniques are often adopted: i) tensor fusion, which merges multiple tensors to transmit them together to reduce the communication startup overhead, and ii) gradient sparsification compression, which transmits only the largest gradient elements to reduce communication traffic. Recent studies focus on combining tensor fusion and sparsification: a straightforward way is to perform sparsification on each tensor before fusion (per-tensor sparsification), but it incurs significant sparsification overhead; alternatively, state-ofthe-art works first merge multiple tensors (i.e., fusion) and then perform sparsification on them to reduce sparsification overhead (multi-tensor sparsification). However, we observe that multi-tensor sparsification causes many tensors to be missing, leading to a significant decrease in convergence accuracy. In this paper, we revisit per-tensor sparsification for no tensor missing and propose a new selective sparsification mechanism tailored for fused tensors. This mechanism strategically selects a subset of fused tensors: the selected ones undergo individual sparsification, while the others remain unsparsified. This design achieves both low sparsification overhead (via selective sparsification) and high convergence performance (without tensor missing), which we term SSFusion. On top of SSFusion, we propose an efficient sparsification offloading scheme that offloads GPU-based gradient sparsification to the CPU to further speed up sparsification, and an interleaved communication scheme that improves communication efficiency through fusion separation. Evaluations on cloud and local clusters show that SSFusion improves the training throughput by 27.5%−106.7%\mathbf{2 7. 5 \% - 1 0 6. 7 \%} over stateof-the-art solutions, while maintaining approximately the same convergence accuracy as the non-sparsification baseline.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖