A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPU
Jiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang, Wei Lin, Wei Zhang
Abstract
While Deep Neural Networks (DNNs) are becoming increasingly popular, there is a growing trend to accelerate the DNN applications on hardware platforms like GPUs, FPGAs, etc., to gain higher performance and efficiency. However, it is time-consuming to tune the performance for such platforms due to the large design space and the expensive cost to evaluate each design point. Although many tuning algorithms, such as XGBoost tuner and genetic algorithm (GA) tuner, have been proposed to guide the design space exploring process in the previous work, the timing issue still remains a critical problem. In this work, we propose a novel auto-tuning framework to optimize the DNN operator design on GPU by leveraging the tuning history efficiently in different scenarios. Our experiments show that we can achieve superior performance than the state-of-the-art work, such as auto-tuning framework TVM and the handcraft optimized library cuDNN, while reducing the searching time by 8.96x and 4.58x comparing with XGBoost tuner and GA tuner in TVM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 13262604-ebef-47ad-8ddf-a64450b2f11fCited by top-tier papers6
- NAAS: Neural Accelerator Architecture SearchYujun Lin, Mengtian Yang, Song HanDAC 2021 · 60 citations
- autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesDu Wu, Jintao Meng, Wenxi Zhu, Minwen Deng et al.SC 2024 · 14 citations
- GTuner: tuning DNN computations on GPU via graph attention networkQi Sun, Xinyun Zhang, Hao Geng, Yuxuan Zhao et al.DAC 2022 · 10 citations
- Bootstrapping in-situ workflow auto-tuning via combining performance models of component applicationsTong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding et al.SC 2021 · 9 citations
- Fast and Efficient DNN Deployment via Deep Gaussian Transfer LearningQi Sun, Chen Bai, Tinghuan Chen, Hao Geng et al.ICCV 2021 · 7 citations
Related papers
- ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program ReuseJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2022 · 10 citations
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 27 citations
- DeepBurning-SEG: Generating DNN Accelerators of Segment-Grained Pipeline ArchitectureXuyi Cai, Ying Wang, Xiaohan Ma, Yinhe Han et al.MICRO 2022 · 25 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- Understanding and bridging the gaps in current GNN performance optimizationsKezhao Huang, Jidong Zhai, Zhen Zheng, Youngmin Yi et al.PPoPP 2021 · 87 citations
