Efficient Direct Convolution Using Long SIMD Instructions
Alexandre de Limas Santana, Adrià Armejach, Marc Casas
摘要
This paper demonstrates that state-of-the-art proposals to compute convolutions on architectures with CPUs supporting SIMD instructions deliver poor performance for long SIMD lengths due to frequent cache conflict misses. We first discuss how to adapt the state-of-the-art SIMD direct convolution to architectures using long SIMD instructions and analyze the implications of increasing the SIMD length on the algorithm formulation. Next, we propose two new algorithmic approaches: the Bounded Direct Convolution (BDC), which adapts the amount of computation exposed to mitigate cache misses, and the Multi-Block Direct Convolution (MBDC), which redefines the activation memory layout to improve the memory access pattern. We evaluate BDC, MBDC, the state-of-the-art technique, and a proprietary library on an architecture featuring CPUs with 16,384-bit SIMD registers using ResNet convolutions. Our results show that BDC and MBDC achieve respective speed-ups of 1.44× and 1.28× compared to the state-of-the-art technique for ResNet-101, and 1.83× and 1.63× compared to the proprietary library.
• Theory of computation → Design and analysis of algorithms; • Computer systems organization → Single instruction, multiple data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu 等HPCA 2025 · 被引用 4 次
- HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise ManagementYao Chen, Cheng Gong, Bingsheng HeMICRO 2025 · 被引用 2 次
- SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream RegistersPaul Scheffler, Luca Colagrande, Luca BeniniDAC 2024 · 被引用 3 次
- GCD2: A Globally Optimizing Compiler for Mapping DNNs to Mobile DSPsWei Niu, Jiexiong Guan, Xipeng Shen, Yanzhi Wang 等MICRO 2022 · 被引用 7 次
- WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field ConvolutionWan Song, Zhou Wei, Rui Wang, Jun Yu 等ICML 2026
