Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning Acceleration
Hang Lu, Liang Chang, Chenglong Li, Zixuan Zhu, Shengjian Lu, Yanhuan Liu, Mingzhe Zhang
Abstract
Along with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity to the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator design called “Bitlet” to maximally exploit the bit-level sparsity. Apart from existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81 × /21 × energy efficiency improvement for training/inference over recent high performance GPUs; (2) up to 15 × /8 × higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (float32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) highly configurable justified by ablation and sensitivity studies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f655882a-4f43-44ab-a048-6ac127efbc05Cited by top-tier papers12
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer et al.HPCA 2024 · 46 citations
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 25 citations
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons et al.HPCA 2025 · 15 citations
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and RepetitivenessHuizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long et al.MICRO 2025 · 10 citations
- Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural NetworksChiyue Wei, Bowen Duan, Cong Guo, Jingyang Zhang et al.ISCA 2025 · 9 citations
Related papers
- BitRed: Taming Non-Uniform Bit-Level Sparsity with a Programmable RISC-V ISA for DNN AccelerationYanhuan Liu, Wenming Li, Kunming Zhang, Yuqun Liu et al.ASPLOS 2026 · 1 citation
- BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionSeunghyun Lee, Dongho Ha, Sungbin Kim, Sungwoo Kim et al.MICRO 2025 · 2 citations
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li et al.DAC 2025
- Bit-slice Architecture for DNN Acceleration with Slice-level Sparsity Enhancement and ExploitationInsu Choi, Young-Seo Yoon, Joon-Sung YangHPCA 2025 · 2 citations
- Bit-Parallel Vector Composability for Neural AccelerationSoroush Ghodrati, Hardik Sharma, Cliff Young, Nam Sung Kim et al.DAC 2020 · 19 citations
