ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN Training
Zhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng, Chenxuan Yao, Dan Feng
2024年份
5被引次数
摘要
Data-parallel deep neural networks (DNN) training systems deployed across nodes have been widely used in various domains, while the system performance is often bottlenecked by the communication overhead among workers for synchronizing gradients. Top-k sparsification compression is the de facto approach to alleviate the communication bottleneck, which truncates the gradient to its largest k elements before sending it to other nodes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SAFusion: Efficient Tensor Fusion with Sparsification Ahead for High-Performance Distributed DNN TrainingZhangqiang Ming, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou 等HPDC 2025 · 被引用 1 次
- Rethinking gradient sparsification as total error minimizationAtal Narayan Sahu, Aritra Dutta, Ahmed M. Abdelmoniem, Trambak Banerjee 等NeurIPS 2021 · 被引用 85 次
- SwitchTop-k: Scaling Top-k Compression on Programmable SwitchesYijun Li, Jiawei Huang, Jingling Liu, Zhaoyi Li 等KDD 2025
- SSFusion: Tensor Fusion with Selective Sparsification for Efficient Distributed DNN TrainingZhangqiang Ming, Rui Wang, Yuchong Hu, Yuanhao Shu 等ICDE 2026 · 被引用 1 次
- DRAGONN: Distributed Randomized Approximate Gradients of Neural NetworksZhuang Wang, Zhaozhuo Xu, Xinyu Crystal Wu, Anshumali Shrivastava 等ICML 2022 · 被引用 10 次
