SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN Training
Zhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu, Wenxiang Zhou, Xinjue Zheng, Dan Feng
摘要
Distributed deep neural networks (DNN) training systems deployed across multiple workers have been widely used to accelerate the training of large models and datasets, while the communication overhead for synchronizing data (i.e., gradient tensor) among workers often becomes a performance bottleneck. To improve communication efficiency, two key techniques are often adopted: i) tensor fusion, which merges multiple tensors to transmit them together to reduce the communication startup overhead, and ii) gradient sparsification compression, which transmits only the largest gradient elements to reduce communication traffic. Recent studies focus on combining tensor fusion and sparsification: a straightforward way is to perform sparsification on each tensor before fusion (per-tensor sparsification), but it incurs significant sparsification overhead; alternatively, state-ofthe-art works first merge multiple tensors (i.e., fusion) and then perform sparsification on them to reduce sparsification overhead (multi-tensor sparsification). However, we observe that multi-tensor sparsification causes many tensors to be missing, leading to a significant decrease in convergence accuracy. In this paper, we revisit per-tensor sparsification for no tensor missing and propose a new selective sparsification mechanism tailored for fused tensors. This mechanism strategically selects a subset of fused tensors: the selected ones undergo individual sparsification, while the others remain unsparsified. This design achieves both low sparsification overhead (via selective sparsification) and high convergence performance (without tensor missing), which we term SSFusion. On top of SSFusion, we propose an efficient sparsification offloading scheme that offloads GPU-based gradient sparsification to the CPU to further speed up sparsification, and an interleaved communication scheme that improves communication efficiency through fusion separation. Evaluations on cloud and local clusters show that SSFusion improves the training throughput by over stateof-the-art solutions, while maintaining approximately the same convergence accuracy as the non-sparsification baseline.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingZhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou 等HPDC 2025 · 被引用 1 次
- ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN TrainingZhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng 等HPDC 2024 · 被引用 5 次
- Communication-Efficient Distributed Deep Learning with Merged Gradient Sparsification on GPUsShaohuai Shi, Qiang Wang, Xiaowen Chu, Bo Li 等INFOCOM 2020 · 被引用 66 次
- Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksYunzhuo Liu, Bo Jiang, Shizhen Zhao, Tao Lin 等INFOCOM 2023 · 被引用 4 次
- ZEN: Empowering Distributed Training with Sparsity-driven Data SynchronizationZhuang Wang, Zhaozhuo Xu, Jingyi Xi, Yuke Wang 等OSDI 2025 · 被引用 3 次
