FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural Networks
Jianxun Yang, Zhao Zhang, Zhuangzhi Liu, Jing Zhou, Leibo Liu, Shaojun Wei, Shouyi Yin
Abstract
Bit-serial computation has been a prevailing convolution method to accelerate varying-precision DNNs by slicing a multi-bit data into multiple 1-bit data and transforming a multiplication into multiple additions, where additions of zero bits are ineffectual, while additions of non-zero bits are repetitive since multiple kernels are quite possible to possess non-zero bits at the same kernel positions. Previous bit-serial accelerators only remove ineffectual additions by skipping computation of zero bits, however, repetitive additions are unable to be eliminated since they compute convolution of each kernel independently. In this work, we propose fused kernel convolution algorithm to eliminate both ineffectual and repetitive additions in bit-serial computation by exploiting bit repetition and bit sparsity in weights, for both convolutional and fully-connected layers. It unifies convolutions of multiple kernels into convolution of one fused kernel by firstly grouping additions into different patterns and secondly reconstructing convolution results, minimizing addition count. Meantime, the memory accesses of activations and partial sums are decreased due to less convolution count. Then a fused kernel convolution based accelerator, FuseKNA, is designed with compact compute logic, which fully exploits value sparsity of activations and bit sparsity of weights. Benchmarked with a set of mainstream DNNs, FuseKNA improves performance by , and , energy efficiency by , and over state-of-the-art Stripes, Pragmatic and Bit-Tactical.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b7b1f403-dbd3-4386-9a0a-834888afca04Cited by top-tier papers5
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng et al.ISCA 2025 · 17 citations
- BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point CloudsMeng Han, Liang Wang, Limin Xiao, Hao Zhang et al.ISCA 2024 · 14 citations
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long et al.MICRO 2025 · 10 citations
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue et al.HPCA 2026 · 2 citations
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng et al.HPCA 2025 · 2 citations
Related papers
- Bit-Serial Cache: Exploiting Input Bit Vector Repetition to Accelerate Bit-Serial InferenceYun-Chen Lo, Ren-Shuo LiuDAC 2023 · 5 citations
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li et al.DAC 2025
- AdaS: A Fast and Energy-Efficient CNN Accelerator Exploiting Bit-SparsityXiaolong Lin, Gang Li, Zizhao Liu, Yadong Liu et al.DAC 2023 · 11 citations
- BitPruner: Network Pruning for Bit-serial AcceleratorsXiandong Zhao, Ying Wang, Cheng Liu, Cong Shi et al.DAC 2020 · 29 citations
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 25 citations
