COMPASS: SRAM-Based Computing-in-Memory SNN Accelerator with Adaptive Spike Speculation
Zongwu Wang, Fangxin Liu, Ning Yang, Shiyuan Huang, Haomin Li, Li Jiang
Abstract
Brain-inspired spiking neural networks (SNNs) are considered energy-efficient alternatives to conventional deep neural networks (DNNs). By adopting event-driven information processing, SNNs can significantly reduce the computational demands associated with DNNs, while still achieving comparable performance. However, current SNNs primarily prioritize high accuracy by constructing complex neuron models that generate sparse spikes. Unfortunately, this approach results in low energy efficiency and high latency, posing a significant challenge for deploying SNNs at the edge. Furthermore, the dominant computation in SNNs, which involves spike-wise Accumulate-Compare operations, is well-suited for Computing-in-Memory (CIM) architectures. However, exploiting high parallel processing and spike sparsity in CIM-based SNN accelerators is challenging due to the irregularity and time dependency of spikes. To address these limitations, the paper proposes COMPASS, a SRAM-based CIM architecture for efficient SNNs. We first introduce an efficient method to exploit irregular sparsity for both input spikes (explicit) and output spikes (implicit). This is achieved through a speculation mechanism that exploit dynamic spike patterns, enabling lean hardware for sparsity utilization. Additionally, the CIM architecture is carefully modified to facilitate dynamic spike pattern generation and exploitation with minimal overhead. Moreover, we design an adaptive dataflow with temporal spike representation tailored for input/output spikes, reducing memory footprint and enabling parallel execution. Comprehensive evaluation results demonstrate that COMPASS can achieve 26.7x end-to-end speedup over recent SNN accelerators hardware implementation with up to 386.7x less energy per inference.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ca16786b-0831-443d-82f1-53e5480221bfCited by top-tier papers1
Ask how each one uses itRelated papers
- Energy-efficient SNN Architecture using 3nm FinFET Multiport SRAM-based CIM with Online LearningLucas Huijbregts, Hsiao-Hsuan Liu, Paul Detterer, Said Hamdioui et al.DAC 2024 · 8 citations
- Resource Constrained Model Compression via Minimax Optimization for Spiking Neural NetworksJue Chen, Huan Yuan, Jianchao Tan, Bin Chen et al.ACM MM 2023 · 5 citations
- Input-Aware Dynamic Timestep Spiking Neural Networks for Efficient In-Memory ComputingYuhang Li, Abhishek Moitra, Tamar Geller, Priyadarshini PandaDAC 2023 · 17 citations
- LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural NetworksRuokai Yin, Youngeun Kim, Di Wu, Priyadarshini PandaMICRO 2024 · 19 citations
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin et al.ISCA 2020 · 122 citations
