HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise Management
Yao Chen, Cheng Gong, Bingsheng He
Abstract
Quantized Deep Neural Networks (DNNs) have progressed to utilize sub-8-bit data types, achieving notable reductions in both memory usage and computational expenses.Nevertheless, the efficient execution of sub-8-bit convolution operations remains insufficiently optimized, especially in the context of Single Instruction Multiple Data (SIMD) architectures.This paper introduces HiPACK, which takes the packing for efficient convolution as foundation to reduce the required number of operations, and address the challenges of SIMD incompatibility in sub-8-bit direct convolution by decoupling the unpacking phase from multiplication operations, optimizing the caching of intermediate data, determining the ideal segmentation, and employing a dual interleaved register (DIR) mechanism to facilitate parallel multiplication through SIMD.Collectively, these optimizations significantly enhance computational efficiency and diminish operational overhead.Evaluation on the BCM2711 ARM processor reveals speedups of up to 4.6× compared to floating-point calculations and 1.7× when measured against state-of-the-art solutions, thus accelerating model inference.Our implementation is open-sourced and available at: https://github.com/Xtra-Computing/HIPACK.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a04d6d6d-5690-4d7a-a01d-025e2342aa22Cited by top-tier papers1
Ask how each one uses itRelated papers
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong et al.SC 2023 · 6 citations
- PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsRuokai Yin, Yuhang Li, Priyadarshini PandaDAC 2025 · 1 citation
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó et al.HPCA 2023 · 31 citations
- Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP PackingJingwei Zhang, Meng Zhang, Xinye Cao, Guoqing LiDAC 2023 · 9 citations
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 18 citations
