Ristretto: An Atomized Processing Architecture for Sparsity-Condensed Stream Flow in CNN
Gang Li, Weixiang Xu, Zhuoran Song, Naifeng Jing, Jian Cheng, Xiaoyao Liang
摘要
Low-precision quantization and sparsity have been widely explored in CNN acceleration due to their effectiveness in reducing computational complexity and memory requirements. However, to support variable numerical precision and sparse computation, prior accelerators design flexible multipliers or sparse dataflow separately. A uniform solution that simultaneously exploits mixed-precision and dual-sided irregular sparsity for CNN acceleration is still lacking. Through an in-depth review of existing precision-scalable and sparse accelerators, we observe that a direct combination of low-level multipliers and high-level sparse dataflow from both sides is challenging due to their orthogonal design spaces. To this end, in this paper, we propose condensed streaming computation. By representing non-zero weights and activations as atomized streams, the low-level mixed-precision multiplication and high-level sparse convolution can be unified into a shared dataflow through hierarchical data reuse. Based on the condensed streaming computation, we propose Ristretto, an atomized architecture that exploits both mixed-precision and dual-sided irregular sparsity for CNN inference. We implement Ristretto in a 28nm technology node. Extensive evaluations show that Ristretto consistently outperforms three state-of-the-art CNN accelerators, including Bit Fusion, Laconic, and SparTen, in terms of performance and energy efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper9
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer 等HPCA 2024 · 被引用 46 次
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue 等MICRO 2024 · 被引用 31 次
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long 等MICRO 2025 · 被引用 10 次
- Misam: Machine Learning Assisted Dataflow Selection in Accelerators for Sparse Matrix MultiplicationSanjali Yadav, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 被引用 6 次
- Pipirima: Predicting Patterns in Sparsity to Accelerate Matrix AlgebraUbaid Bakhtiar, Donghyeon Joo, Bahar AsgariDAC 2025 · 被引用 6 次
相关 Paper
- AdaS: A Fast and Energy-Efficient CNN Accelerator Exploiting Bit-SparsityXiaolong Lin, Gang Li, Zizhao Liu, Yadong Liu 等DAC 2023 · 被引用 11 次
- FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural NetworksJianxun Yang, Zhao Zhang, Zhuangzhi Liu, Jing Zhou 等HPCA 2021 · 被引用 20 次
- CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric FiltersJiajun Li, Ahmed Louri, Avinash Karanth, Razvan C. BunescuHPCA 2021 · 被引用 9 次
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu 等MICRO 2021 · 被引用 54 次
- PENNI: Pruned Kernel Sharing for Efficient CNN InferenceShiyu Li, Edward Hanson, Hai Li, Yiran ChenICML 2020 · 被引用 23 次
