Bit-Serial Cache: Exploiting Input Bit Vector Repetition to Accelerate Bit-Serial Inference
Yun-Chen Lo, Ren-Shuo Liu
摘要
Bit-serial computation has demonstrated superiority in processing precision-varying DNNs by slicing multi-bit vectors into multiple single-bit vectors and computing the inner product using multiple steps of shift-and-adds. In this paper, we identify that performing real-world DNNs inference with bit-serial computation exhibits high input bit vector locality, where up to 85.7% of non-zero input bit vectors, as well as their associated computation, are previously-seen and previously-done ones. We propose Bit-Serial Cache to transfer the identified locality into performance and energy efficiency gains. The key design strategy is to store recently-computed partial sums of input bit vectors to a cache and utilize cache accesses to replace redundant computations. In addition to the bit-serial computation architecture, we also present: 1) request clustering and 2) interleaved scheduling, to further enhance the performance and energy efficiency.Our experiments using six popular DNNs (in both 8-b and 4-b) show that Bit-Serial Cache speeds up DNN inference by up to 2.72×, 1.82×, and 4.03×, energy efficiency by 3.19×, 3.29×, and 2.82×, area efficiency by 1.35×, 1.24×, and 2.76× over state-of-the-art Loom, DPRed Loom, and Laconic.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long 等MICRO 2025 · 被引用 10 次
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue 等HPCA 2026 · 被引用 2 次
相关 Paper
- FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural NetworksJianxun Yang, Zhao Zhang, Zhuangzhi Liu, Jing Zhou 等HPCA 2021 · 被引用 20 次
- AdaS: A Fast and Energy-Efficient CNN Accelerator Exploiting Bit-SparsityXiaolong Lin, Gang Li, Zizhao Liu, Yadong Liu 等DAC 2023 · 被引用 11 次
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer 等HPCA 2024 · 被引用 46 次
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li 等DAC 2025
- BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionSeunghyun Lee, Dongho Ha, Sungbin Kim, Sungwoo Kim 等MICRO 2025 · 被引用 2 次
