Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
Chiyue Wei, Cong Guo, Feng Cheng, Shiyu Li, Hao (Frank) Yang, Hai Helen Li, Yiran Chen
摘要
Spiking Neural Networks (SNNs) are highly efficient due to their spike-based activation, which inherently produces bit-sparse computation patterns. Existing hardware implementations of SNNs leverage this sparsity pattern to avoid wasteful zero-value computations, yet this approach fails to fully capitalize on the potential efficiency of SNNs. This study introduces a novel sparsity paradigm called Product Sparsity, which leverages combinatorial similarities within matrix multiplication operations to reuse the inner product result and reduce redundant computations. Product Sparsity significantly enhances sparsity in SNNs without compromising the original computation results compared to traditional bit sparsity methods. For instance, in the SpikeBERT SNN model, Product Sparsity achieves a density of only 1.23% and reduces computation by , compared to bit sparsity, which has a density of 13.19%. To efficiently implement Product Sparsity, we propose Prosperity, an architecture that addresses the challenges of identifying and eliminating redundant computations in real-time. Compared to prior SNN accelerator PTB and the A100 GPU, Prosperity achieves an average speedup of and , respectively, along with energy efficiency improvements of and , respectively. The code for Prosperity is available at https://github.com/dubcyfor3/Prosperity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-Aware Cache CompressionFeng Cheng, Cong Guo, Chiyue Wei, Junyao Zhang 等ISCA 2025 · 被引用 12 次
- Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural NetworksChiyue Wei, Bowen Duan, Cong Guo, Jingyang Zhang 等ISCA 2025 · 被引用 9 次
- Transitive Array: An Efficient GEMM Accelerator with Result ReuseCong Guo, Chiyue Wei, Jiaming Tang, Bowen Duan 等ISCA 2025 · 被引用 8 次
- EVA: Accelerating LLM Decoding via an Efficient Vector Quantization ArchitectureBowen Duan, Cong Guo, Chiyue Wei, Haoxuan Shan 等ISCA 2026 · 被引用 2 次
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language ModelsChiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan 等HPCA 2026 · 被引用 2 次
它引用的顶会 Paper26
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang 等NeurIPS 2021 · 被引用 857 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan 等NeurIPS 2023 · 被引用 368 次
- Differentiable Spike: Rethinking Gradient-Descent for Training Spiking Neural NetworksYuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng 等NeurIPS 2021 · 被引用 288 次
相关 Paper
- Spik4lite: Refactoring Neuromorphic Sparsity for Efficient Spiking Neural Networks on Commodity Edge DevicesYongzhi She, Qihua Zhou, Yuhao Wang, Yaodong Huang 等ICML 2026
- LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural NetworksRuokai Yin, Youngeun Kim, Di Wu, Priyadarshini PandaMICRO 2024 · 被引用 19 次
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin 等ISCA 2020 · 被引用 122 次
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 被引用 25 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
