Lune

MICRO2025顶会

HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise Management

Yao Chen, Cheng Gong, Bingsheng He

2025年份
2被引次数
1顶会引用

摘要

Quantized Deep Neural Networks (DNNs) have progressed to utilize sub-8-bit data types, achieving notable reductions in both memory usage and computational expenses.Nevertheless, the efficient execution of sub-8-bit convolution operations remains insufficiently optimized, especially in the context of Single Instruction Multiple Data (SIMD) architectures.This paper introduces HiPACK, which takes the packing for efficient convolution as foundation to reduce the required number of operations, and address the challenges of SIMD incompatibility in sub-8-bit direct convolution by decoupling the unpacking phase from multiplication operations, optimizing the caching of intermediate data, determining the ideal segmentation, and employing a dual interleaved register (DIR) mechanism to facilitate parallel multiplication through SIMD.Collectively, these optimizations significantly enhance computational efficiency and diminish operational overhead.Evaluation on the BCM2711 ARM processor reveals speedups of up to 4.6× compared to floating-point calculations and 1.7× when measured against state-of-the-art solutions, thus accelerating model inference.Our implementation is open-sourced and available at: https://github.com/Xtra-Computing/HIPACK.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖