Romou: rapidly generate high-performance tensor kernels for mobile GPUs
Rendong Liang, Ting Cao, Jicheng Wen, Manni Wang, Yang Wang, Jianhua Zou, Yunxin Liu
摘要
Mobile GPU, as a ubiquitous and powerful accelerator, plays an important role in accelerating on-device DNN (Deep Neural Network) inference. The frequent-upgrade and diversity of mobile GPUs require automatic kernel generation to empower fast DNN deployment. However, current generated kernels have poor performance.
The goal of this paper is to rapidly generate high-performance kernels for diverse mobile GPUs. The major challenges are (1) it is unclear about what is the optimal kernel due to the lack of hardware knowledge; (2) how to rapidly generate it from a large space of candidates. For the first challenge, we propose a crossplatform profiling tool, the first to disclose and quantify mobile GPU architecture. The result demystifies the hardware bottleneck, and also directs the solution for the second challenge by exposing the unique high-performance hardware feature, identifying inefficient kernels against hardware constraints, and specifying performance bound for kernels.
Directed by that, we propose a mobile-GPU-specific kernel compiler Romou. It supports the unique hardware feature in kernel implementation, and prunes inefficient ones against hardware resources. Romou can thus rapidly generate high-performance kernels. Compared to the state-of-the-art generated kernels, it achieves up-to 14.7× speedup on average for convolution. Up-to 99% search space is pruned. The performance is even up-to 1.2× faster on average than the state-of-the-art hand-optimized implementation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ROLLER: Fast and Efficient Tensor Compilation for Deep LearningHongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke 等OSDI 2022 · 被引用 84 次
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang 等MobiCom 2023 · 被引用 29 次
- MobiDepth: real-time depth estimation using on-device dual camerasJinrui Zhang, Huan Yang, Ju Ren, Deyu Zhang 等MobiCom 2022 · 被引用 24 次
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on MobileWei Niu, Md. Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan 等ASPLOS 2024 · 被引用 11 次
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu 等ASPLOS 2026
它引用的顶会 Paper4
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 被引用 90 次
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 被引用 71 次
相关 Paper
- RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech RecognitionPeiyan Dong, Siyue Wang, Wei Niu, Chengming Zhang 等DAC 2020 · 被引用 50 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesWei Niu, Mengshu Sun, Zhengang Li, Jou-An Chen 等AAAI 2021 · 被引用 14 次
- DeepCache: Revisiting Cache Side-Channel Attacks in Deep Neural Networks ExecutablesZhibo Liu, Yuanyuan Yuan, Yanzuo Chen, Sihang Hu 等CCS 2024 · 被引用 3 次
- DyCL: Dynamic Neural Network Compilation Via Program Rewriting and Graph OptimizationSimin Chen, Shiyi Wei, Cong Liu, Wei YangISSTA 2023 · 被引用 11 次
