Bit-Parallel Vector Composability for Neural Acceleration
Soroush Ghodrati, Hardik Sharma, Cliff Young, Nam Sung Kim, Hadi Esmaeilzadeh
Abstract
Conventional neural accelerators rely on isolated self-sufficient functional units that perform an atomic operation while communicating the results through an operand delivery-aggregation logic. Each single unit processes all the bits of their operands atomically and produce all the bits of the results in isolation. This paper explores a different design style, where each unit is only responsible for a slice of the bit-level operations to interleave and combine the benefits of bit-level parallelism with the abundant data-level parallelism in deep neural networks. A dynamic collection of these units cooperate at runtime to generate bits of the results, collectively. Such cooperation requires extracting new grouping between the bits, which is only possible if the operands and operations are vectorizable. The abundance of Data-Level Parallelism and mostly repeated execution patterns, provides a unique opportunity to define and leverage this new dimension of Bit-Parallel Vector Composability. This design intersperses bit parallelism within data-level parallelism and dynamically interweaves the two together. As such, the building block of our neural accelerator is a Composable Vector Unit that is a collection of Narrower-Bitwidth Vector Engines, which are dynamically composed or decomposed at the bit granularity. Using six diverse CNN and LSTM deep networks, we evaluate this design style across four design points: with and without algorithmic bitwidth heterogeneity and with and without availability of a high-bandwidth off-chip memory. Across these four design points, Bit-Parallel Vector Composability brings (1.4× to 3.5×) speedup and (1.1× to 2.7×) energy reduction. We also comprehensively compare our design style to the Nvidia's RTX 2080 TI GPU, which also supports INT-4 execution. The benefits range between 28.0× and 33.7× improvement in Performance-per-Watt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ff67b2b-ed5b-4e70-b9bb-7a02bf0346b8Cited by top-tier papers5
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer et al.MICRO 2020 · 120 citations
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationAmir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu KangMICRO 2022 · 47 citations
- Tandem Processor: Grappling with Emerging Operators in Neural NetworksSoroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra et al.ASPLOS 2024 · 21 citations
- In-Storage Domain-Specific Acceleration for Serverless ComputingRohan Mahapatra, Soroush Ghodrati, Byung Hoon Ahn, Sean Kinzer et al.ASPLOS 2024 · 7 citations
- DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel ProcessingYansong Xu, Dongxu Lyu, Zhenyu Li, Yuzhou Chen et al.DAC 2024 · 5 citations
Related papers
- Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationHang Lu, Liang Chang, Chenglong Li, Zixuan Zhu et al.MICRO 2021 · 54 citations
- BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionSeunghyun Lee, Dongho Ha, Sungbin Kim, Sungwoo Kim et al.MICRO 2025 · 2 citations
- BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern PruningGang Wang, Siqi Cai, Zhenyu Li, Wenjie Li et al.DAC 2025
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer et al.HPCA 2024 · 46 citations
