SC2023Top-tier venue
Optimizing Direct Convolutions on ARM Multi-Cores
Pengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang, Peng Zhang, Tao Tang, Zheng Wang
Abstract
Convolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8106054e-9460-491a-bf42-fb39e86b5f76Cited by top-tier papers2
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsHaoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu et al.AAAI 2026
Builds on5
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev et al.ASPLOS 2021 · 52 citations
- Efficient Direct Convolution Using Long SIMD InstructionsAlexandre de Limas Santana, Adrià Armejach, Marc CasasPPoPP 2023 · 14 citations
- Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloadsEvangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman et al.SC 2021 · 2 citations
Related papers
- HiPACK: Efficient Sub-8-Bit Direct Convolution with SIMD and Bitwise ManagementYao Chen, Cheng Gong, Bingsheng HeMICRO 2025 · 2 citations
- High Performance Depthwise and Pointwise Convolutions on Mobile DevicesPengfei Zhang, Eric Lo, Baotong LuAAAI 2020 · 54 citations
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 27 citations
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using F_2Keren Zhou, Mario Lezcano Casado, Adam P. Goucher, Akhmed Rakhmati et al.ASPLOS 2026 · 3 citations
