High Performance Depthwise and Pointwise Convolutions on Mobile Devices
Pengfei Zhang, Eric Lo, Baotong Lu
Abstract
Lightweight convolutional neural networks (e.g., Mo-bileNets) are specifically designed to carry out inference directly on mobile devices. Among the various lightweight models, depthwise convolution (DWConv) and pointwise convolution (PWConv) are their key operations. In this paper, we observe that the existing implementations of DW-Conv and PWConv are not well utilizing the ARM processors in the mobile devices, and exhibit lots of cache misses under multi-core and poor data reuse at register level. We propose techniques to re-optimize the implementations of DWConv and PWConv based on ARM architecture. Experimental results show that our implementation can respectively achieve a speedup of up to 5.5× and 2.1× against TVM (Chen et al. 2018) on DWConv and PWConv.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1e71aa0-b14d-4030-9247-c40624326c3fCited by top-tier papers6
- Multi-Frequency Representation Enhancement with Privilege Information for Video Super-ResolutionFei Li, Linfeng Zhang, Zikun Liu, Juan Lei et al.ICCV 2023 · 24 citations
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 18 citations
- XiNet: Efficient Neural Networks for tinyMLAlberto Ancilotto, Francesco Paissan, Elisabetta FarellaICCV 2023 · 18 citations
- Reverse Convolution and its Applications to Image RestorationXuhong Huang, Shiqi Liu, Kai Zhang, Ying Tai et al.ICCV 2025 · 13 citations
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon et al.NeurIPS 2024 · 10 citations
Related papers
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
- MobileODE: An Extra Lightweight NetworkLe Yu, Jun Wu, Bo Gou, Xiangde Min et al.NeurIPS 2025 · 8 citations
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong et al.SC 2023 · 6 citations
- On The Efficiency of Sparse-Tiled Tensor Graph Processing For Low Memory UsageAntonio Cipolletta, Andrea CalimeraDAC 2021 · 6 citations
- Dynamic Convolution: Attention Over Convolution KernelsYinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen et al.CVPR 2020
