GCD2: A Globally Optimizing Compiler for Mapping DNNs to Mobile DSPs
Wei Niu, Jiexiong Guan, Xipeng Shen, Yanzhi Wang, Gagan Agrawal, Bin Ren
Abstract
More specialized chips are exploiting available high transistor density to expose parallelism at a large scale with more intricate instruction sets. This paper reports on a compilation system GCD2, developed to support complex Deep Neural Network (DNN) workloads on mobile DSP chips. We observe several challenges in fully exploiting this architecture, related to SIMD width, more complex SIMD/vector instructions, and VLIW pipeline with the notion of soft dependencies. GCD2comprises the following contributions: 1) development of matrix layout formats that support the use of different novel SIMD instructions, 2) formulation and solution of a global optimization problem related to choosing the best instruction (and associated layout) for implementation of each operator in a complete DNN, and 3) SDA, an algorithm for packing instructions with consideration for soft dependencies. These solutions are incorporated in a complete compilation system that is extensively evaluated against other systems using 10 large DNN models. Evaluation results show that GCD2outperforms two product-level state-of-the-art end-to-end DNN execution frameworks (TFLite and Qualcomm SNPE) that support mobile DSPs by up to speedup, and outperforms three established compilers (Halide, TVM, and RAKE) by up to and speedup, respectively. GCD2is also unique in supporting, real-time execution of certain DNNs, while its implementation enables two major DNNs to execute on a mobile DSP for the first time.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0decdd41-336c-4a88-a3f3-6c6e83a88c7dCited by top-tier papers3
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu et al.ASPLOS 2025 · 38 citations
- NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingYitu Wang, Shiyu Li, Qilin Zheng, Linghao Song et al.ISCA 2024 · 26 citations
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
Related papers
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang et al.ASPLOS 2024 · 14 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- SoD2: Statically Optimizing Dynamic Deep Neural Network ExecutionWei Niu, Gagan Agrawal, Bin RenASPLOS 2024 · 6 citations
- Romou: rapidly generate high-performance tensor kernels for mobile GPUsRendong Liang, Ting Cao, Jicheng Wen, Manni Wang et al.MobiCom 2022 · 13 citations
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
