High Performance Depthwise and Pointwise Convolutions on Mobile Devices
Pengfei Zhang, Eric Lo, Baotong Lu
摘要
Lightweight convolutional neural networks (e.g., Mo-bileNets) are specifically designed to carry out inference directly on mobile devices. Among the various lightweight models, depthwise convolution (DWConv) and pointwise convolution (PWConv) are their key operations. In this paper, we observe that the existing implementations of DW-Conv and PWConv are not well utilizing the ARM processors in the mobile devices, and exhibit lots of cache misses under multi-core and poor data reuse at register level. We propose techniques to re-optimize the implementations of DWConv and PWConv based on ARM architecture. Experimental results show that our implementation can respectively achieve a speedup of up to 5.5× and 2.1× against TVM (Chen et al. 2018) on DWConv and PWConv.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Multi-Frequency Representation Enhancement with Privilege Information for Video Super-ResolutionFei Li, Linfeng Zhang, Zikun Liu, Juan Lei 等ICCV 2023 · 被引用 24 次
- Hybrid Neural Networks for On-Device Directional HearingAnran Wang, Maruchi Kim, Hao Zhang, Shyamnath GollakotaAAAI 2022 · 被引用 18 次
- XiNet: Efficient Neural Networks for tinyMLAlberto Ancilotto, Francesco Paissan, Elisabetta FarellaICCV 2023 · 被引用 18 次
- Reverse Convolution and its Applications to Image RestorationXuhong Huang, Shiqi Liu, Kai Zhang, Ying Tai 等ICCV 2025 · 被引用 13 次
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
相关 Paper
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin 等AAAI 2020 · 被引用 201 次
- MobileODE: An Extra Lightweight NetworkLe Yu, Jun Wu, Bo Gou, Xiangde Min 等NeurIPS 2025 · 被引用 8 次
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong 等SC 2023 · 被引用 6 次
- On The Efficiency of Sparse-Tiled Tensor Graph Processing For Low Memory UsageAntonio Cipolletta, Andrea CalimeraDAC 2021 · 被引用 6 次
- Dynamic Convolution: Attention Over Convolution KernelsYinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 等CVPR 2020
