LUTein: Dense-Sparse Bit-Slice Architecture With Radix-4 LUT-Based Slice-Tensor Processing Units
Dongseok Im, Hoi-Jun Yoo
Abstract
Bit-slice architectures have been developed to support various bit-precision and data sparsity of deep neural networks (DNNs). However, because of low-bit precision and a wide range of sparsity of bit-slice computations, bit-slice architectures are challenging in multiplier-, processing element (PE)-, core-, and software (SW)-level designs. First, data sparsity causes a power trade-off between the Radix numbers of a multiplier, and the previous multipliers cannot take advantage of all sparsity ranges. Second, a bit-slice PE which integrated massive lowbit multiplier-and-accumulate (MAC) units brings about large data transactions compared to a fixed bit-width PE. Third, bitslice core architectures only focus on either dense or sparse data computations, limiting the overall performance of bit-slice computations with a wide sparsity range. Lastly, low-bit bit-slice computations cause massive repetitive instruction fetches across hardware units. To solve the challenges, LUTein is proposed. It exploits the new lookup table (LUT)-based computing method to support the Radix-4 Modified Booth algorithm, achieving low power consumption in all sparsity ranges. Moreover, the slice-tensor PE efficiently processes slice-tensor data by sharing hardware units across the Radix-4 LUT-based MAC units. In addition, the LUTein architecture adopts a systolic datapath with a multi-port buffer to exploit both inter-PE data reuse and slice-level sparsity. Lastly, LUTein's instruction set architecture (ISA) and the hierarchical instruction decoder are introduced to alleviate repetitive instruction fetches. As a result, LUTein outperforms the state-of-the-art bit-slice architecture, Sibia, over 1.34× higher energy-efficiency and 1.78× higher area-efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6880f10a-42fa-4db5-b76a-ad88d85837d5Cited by top-tier papers2
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng et al.ISCA 2025 · 17 citations
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng et al.HPCA 2025 · 2 citations
Related papers
- Bit-slice Architecture for DNN Acceleration with Slice-level Sparsity Enhancement and ExploitationInsu Choi, Young-Seo Yoon, Joon-Sung YangHPCA 2025 · 2 citations
- Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity ExploitationDongseok Im, Gwangtae Park, Zhiyong Li, Junha Ryu et al.HPCA 2023 · 26 citations
- BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionSeunghyun Lee, Dongho Ha, Sungbin Kim, Sungwoo Kim et al.MICRO 2025 · 2 citations
- Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice SparsityDongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong et al.HPCA 2025 · 5 citations
- Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network AccelerationAkshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Srinivasa, Tarun Makesh Chandran et al.MICRO 2020 · 50 citations
