ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic Computing
Kang You, Chen Nie, Lee Jun Yan, Ziling Wei, Cheng Zou, Zekai Xu, Yu Feng, Honglan Jiang, Zhezhi He
摘要
Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A key temporal property of SNNs, elastic inference, allows outputs to emerge progressively, enabling responses to salient inputs much earlier than full evaluation. However, existing SNN-specific accelerators cannot capitalize on this property. Layer-by-layer designs emit outputs only after all layers are complete, while time-step-by-time-step designs rely on coarse-grained, layer-wise pipelines that require synchronizing all spines/tokens within a layer. This barrier prevents results from being forwarded immediately, delaying the earliest possible response and forfeiting the benefits of elastic inference. To address these challenges, we propose ELSA, a near-SRAM dataflow architecture that realizes true elastic inference through a fine-grained spine/token-wise pipeline and hardware optimizations tailored to SNNs. ELSA forwards each spine/token immediately upon production, forming a continuous streaming pipeline that substantially reduces the latency to the first response. To enhance this lightweight execution, ELSA introduces a bundled address event representation protocol to lower communication traffic of network-on-chip (NoC), and leverages mini-batch spiking Gustavson-product to cut memory access and exploit inherent sparsity. Combined with mapping and scheduling optimizations, ELSA achieves efficient, event-driven computation without compromising accuracy. Experiments show that SNNs can outperform quantized artificial neural networks (QANNs) while maintaining on-par accuracy. For a 4-bit ResNet-50, ELSA achieves 3.4× speedup and higher energy efficiency over the SOTA QANN accelerator (ANT), and speedup and energy efficiency gains over the SOTA SNN accelerator (PAICORE).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural NetworksTong Bu, Wei Fang, Jianhao Ding, Penglin Dai 等ICLR 2022 · 被引用 272 次
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi 等MICRO 2020 · 被引用 223 次
- Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable ArchitectureLiqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo 等MICRO 2021 · 被引用 221 次
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 被引用 158 次
相关 Paper
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin 等ISCA 2020 · 被引用 122 次
- SATO: spiking neural network acceleration via temporal-oriented dataflow and architectureFangxin Liu, Wenbo Zhao, Zongwu Wang, Yongbiao Chen 等DAC 2022 · 被引用 20 次
- LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural NetworksRuokai Yin, Youngeun Kim, Di Wu, Priyadarshini PandaMICRO 2024 · 被引用 19 次
- Stellar: Energy-Efficient and Low-Latency SNN Algorithm and Hardware Co-Design with Spatiotemporal ComputationRuixin Mao, Lin Tang, Xingyu Yuan, Ye Liu 等HPCA 2024 · 被引用 21 次
- HardF-SNN: Hardware-Friendly Quantization for Spiking Neural Networks with Efficient Integer-Arithmetic-Only InferenceHanwen Liu, Kexin Shi, Jieyuan Zhang, Yimeng Shan 等AAAI 2026
