HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise Management
Yao Chen, Cheng Gong, Bingsheng He
摘要
Quantized Deep Neural Networks (DNNs) have progressed to utilize sub-8-bit data types, achieving notable reductions in both memory usage and computational expenses.Nevertheless, the efficient execution of sub-8-bit convolution operations remains insufficiently optimized, especially in the context of Single Instruction Multiple Data (SIMD) architectures.This paper introduces HiPACK, which takes the packing for efficient convolution as foundation to reduce the required number of operations, and address the challenges of SIMD incompatibility in sub-8-bit direct convolution by decoupling the unpacking phase from multiplication operations, optimizing the caching of intermediate data, determining the ideal segmentation, and employing a dual interleaved register (DIR) mechanism to facilitate parallel multiplication through SIMD.Collectively, these optimizations significantly enhance computational efficiency and diminish operational overhead.Evaluation on the BCM2711 ARM processor reveals speedups of up to 4.6× compared to floating-point calculations and 1.7× when measured against state-of-the-art solutions, thus accelerating model inference.Our implementation is open-sourced and available at: https://github.com/Xtra-Computing/HIPACK.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong 等SC 2023 · 被引用 6 次
- PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsRuokai Yin, Yuhang Li, Priyadarshini PandaDAC 2025 · 被引用 1 次
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó 等HPCA 2023 · 被引用 31 次
- Uint-Packing: Multiply Your DNN Accelerator Performance via Unsigned Integer DSP PackingJingwei Zhang, Meng Zhang, Xinye Cao, Guoqing LiDAC 2023 · 被引用 9 次
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 被引用 18 次
