DenSparSA: A Balanced Systolic Array Approach for Dense and Sparse Matrix Multiplication
Ziheng Wang, Ruiqi Sun, Xin He, Tianrui Ma, An Zou
摘要
Numerous studies have proposed hardware architectures to accelerate sparse matrix multiplication, but these approaches often incur substantial area and power overhead, significantly compromising their usage in dense scenarios. On the other hand, systolic arrays deliver high efficiency for dense matrix operations, but their application to sparse matrices remains challenging. An ideal design should process both dense and sparse matrices with high efficiency to satisfy performance and versatility requirements.In this paper, we introduce DenSparSA, a balanced systolic array centralized architecture that can execute sparse matrix computations with minimal overhead to original dense matrix computations. DenSparSA supports both single-side and dual-side unstructured sparse matrix multiplications with high efficiency. At the same time, the additional hardware required for managing sparsity is compact and decoupled from the conventional systolic array, allowing for minimal power overhead when switched back to dense matrix operations via circuit gating. The proposed design is implemented with Nangate 45 nm. Implementation results show that DenSparSA achieves a speedup ranging from to compared to the classic systolic array for sparse workloads, while maintaining relatively low area and power overhead. For dense workloads, the power overhead can be reduced to for BF16 and 5% for FP32. Compared with existing solutions for sparse acceleration, DenSparSA delivers competitive () efficiency in sparse scenarios and better efficiency for dense scenarios, indicating a better balance between both situations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HiT: A Unified Sparsity-Adaptive Architecture for High-Throughput Matrix MultiplicationTingting Xiang, Xiaochen Wang, Miao Yu, Trevor E. CarlsonISCA 2026
- Trapezoid: A Versatile Accelerator for Dense and Sparse Matrix MultiplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2024 · 被引用 38 次
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 被引用 37 次
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun 等DAC 2024 · 被引用 3 次
