Efficient Direct Convolution Using Long SIMD Instructions
Alexandre de Limas Santana, Adrià Armejach, Marc Casas
Abstract
This paper demonstrates that state-of-the-art proposals to compute convolutions on architectures with CPUs supporting SIMD instructions deliver poor performance for long SIMD lengths due to frequent cache conflict misses. We first discuss how to adapt the state-of-the-art SIMD direct convolution to architectures using long SIMD instructions and analyze the implications of increasing the SIMD length on the algorithm formulation. Next, we propose two new algorithmic approaches: the Bounded Direct Convolution (BDC), which adapts the amount of computation exposed to mitigate cache misses, and the Multi-Block Direct Convolution (MBDC), which redefines the activation memory layout to improve the memory access pattern. We evaluate BDC, MBDC, the state-of-the-art technique, and a proprietary library on an architecture featuring CPUs with 16,384-bit SIMD registers using ResNet convolutions. Our results show that BDC and MBDC achieve respective speed-ups of 1.44× and 1.28× compared to the state-of-the-art technique for ResNet-101, and 1.83× and 1.63× compared to the proprietary library.
• Theory of computation → Design and analysis of algorithms; • Computer systems organization → Single instruction, multiple data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu et al.HPCA 2025 · 4 citations
- HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise ManagementYao Chen, Cheng Gong, Bingsheng HeMICRO 2025 · 2 citations
- SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream RegistersPaul Scheffler, Luca Colagrande, Luca BeniniDAC 2024 · 3 citations
- GCD2: A Globally Optimizing Compiler for Mapping DNNs to Mobile DSPsWei Niu, Jiexiong Guan, Xipeng Shen, Yanzhi Wang et al.MICRO 2022 · 7 citations
- WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field ConvolutionWan Song, Zhou Wei, Rui Wang, Jun Yu et al.ICML 2026
