TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training
Mostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad, Gennady Pekhimenko, Jorge Albericio, Andreas Moshovos
Abstract
TensorDash is a hardware level technique for enabling data-parallel MAC units to take advantage of sparsity in their input operand streams. When used to compose a hardware accelerator for deep learning, TensorDash can speedup the training process while also increasing energy efficiency. TensorDash combines a low-cost, sparse input operand interconnect comprising an 8-input multiplexer per multiplier input, with an area efficient hardware scheduler. While the interconnect allows a very limited set of movements per operand, the scheduler can effectively extract sparsity when it is present in the activations, weights or gradients of neural networks. Over a wide set of models covering various applications, TensorDash accelerates the training process by 1.95× while being 1.89× more energy efficient, 1.6× more energy efficient when taking on-chip and off-chip memory accesses into account. While TensorDash works with any datatype, we demonstrate it with both single-precision floating-point units and bfloat16.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d8da819-e0a5-4f4d-a2f6-4fcfd1f266bfCited by top-tier papers15
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- S2TA: Exploiting Structured Sparsity for Energy-Efficient Mobile CNN AccelerationZhi Gang Liu, Paul N. Whatmough, Yuhao Zhu, Matthew MattinaHPCA 2022 · 110 citations
- HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationQingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu et al.ISCA 2021 · 73 citations
- FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingSai Qian Zhang, Bradley McDanel, H. T. KungHPCA 2022 · 68 citations
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
Related papers
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu et al.MICRO 2021 · 54 citations
- FPRaker: A Processing Element For Accelerating Neural Network TrainingOmar Mohamed Awad, Mostafa Mahmoud, Isak Edo, Ali Hadi Zadeh et al.MICRO 2021 · 13 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- SPARK: An Efficient Hybrid Acceleration Architecture with Run-Time Sparsity-Aware Scheduling for TinyML LearningMingxuan Li, Qinzhe Zhi, Yanchi Dong, Le Ye et al.DAC 2024 · 2 citations
- Eureka: Efficient Tensor Cores for One-sided Unstructured Sparsity in DNN InferenceAshish Gondimalla, Mithuna Thottethodi, T. N. VijaykumarMICRO 2023 · 15 citations
