LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
Ruokai Yin, Youngeun Kim, Di Wu, Priyadarshini Panda
摘要
Spiking Neural Networks (SNNs) have gained significant research attention in the last decade due to their potential to drive resource-constrained edge devices. Though existing SNN accelerators offer high efficiency in processing sparse spikes with dense weights, opportunities are less explored in SNNs with sparse weights, i.e., dual-sparsity. In this work, we study the acceleration of dual-sparse SNNs, focusing on their core operation, sparse-matrix-sparse-matrix multiplication (spMspM). We observe that naively running a dual-sparse SNN on existing spMspM accelerators designed for dual-sparse Artificial Neural Networks (ANNs) exhibits sub-optimal efficiency. The main challenge is that processing timesteps, a natural property of SNNs, introduces an extra loop to ANN spMspM, leading to longer latency and more memory traffic. To address the problem, we propose a fully temporal-parallel (FTP) dataflow, which minimizes both data movement across timesteps and the endto-end latency of dual-sparse SNNs. To maximize the efficiency of FTP dataflow, we propose an FTP-friendly spike compression mechanism that efficiently compresses single-bit spikes and ensures contiguous memory access. We further propose an FTPfriendly inner-join circuit that can lower the cost of the expensive prefix-sum circuits with almost no throughput penalty. All the above techniques for FTP dataflow are encapsulated in LoAS, a Low-latency inference Accelerator for dual-sparse SNNs. With FTP dataflow, compression, and inner-join, running dual-sparse SNN workloads on LoAS demonstrates significant speedup (up to 8.51×) and energy reduction (up to 3.68×) compared to running it on prior dual-sparse accelerators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Prosperity: Accelerating Spiking Neural Networks via Product SparsityChiyue Wei, Cong Guo, Feng Cheng, Shiyu Li 等HPCA 2025 · 被引用 14 次
- Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-constrained PruningBoxun Xu, Yuxuan Yin, Vikram Iyer, Peng LiISCA 2025 · 被引用 4 次
- ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic ComputingKang You, Chen Nie, Lee Jun Yan, Ziling Wei 等ISCA 2026
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao 等CVPR 2025
它引用的顶会 Paper19
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang 等NeurIPS 2021 · 被引用 857 次
- Going Deeper With Directly-Trained Larger Spiking Neural NetworksHanle Zheng, Yujie Wu, Lei Deng, Yifan Hu 等AAAI 2021 · 被引用 694 次
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object DetectionSei Joon Kim, Seongsik Park, Byunggook Na, Sungroh YoonAAAI 2020 · 被引用 512 次
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan 等NeurIPS 2023 · 被引用 368 次
相关 Paper
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin 等ISCA 2020 · 被引用 122 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
- SATO: spiking neural network acceleration via temporal-oriented dataflow and architectureFangxin Liu, Wenbo Zhao, Zongwu Wang, Yongbiao Chen 等DAC 2022 · 被引用 20 次
- COMPASS: SRAM-Based Computing-in-Memory SNN Accelerator with Adaptive Spike SpeculationZongwu Wang, Fangxin Liu, Ning Yang, Shiyuan Huang 等MICRO 2024 · 被引用 11 次
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao 等ACM MM 2024 · 被引用 23 次
