A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPU
Jiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang, Wei Lin, Wei Zhang
摘要
While Deep Neural Networks (DNNs) are becoming increasingly popular, there is a growing trend to accelerate the DNN applications on hardware platforms like GPUs, FPGAs, etc., to gain higher performance and efficiency. However, it is time-consuming to tune the performance for such platforms due to the large design space and the expensive cost to evaluate each design point. Although many tuning algorithms, such as XGBoost tuner and genetic algorithm (GA) tuner, have been proposed to guide the design space exploring process in the previous work, the timing issue still remains a critical problem. In this work, we propose a novel auto-tuning framework to optimize the DNN operator design on GPU by leveraging the tuning history efficiently in different scenarios. Our experiments show that we can achieve superior performance than the state-of-the-art work, such as auto-tuning framework TVM and the handcraft optimized library cuDNN, while reducing the searching time by 8.96x and 4.58x comparing with XGBoost tuner and GA tuner in TVM.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- NAAS: Neural Accelerator Architecture SearchYujun Lin, Mengtian Yang, Song HanDAC 2021 · 被引用 60 次
- autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesDu Wu, Jintao Meng, Wenxi Zhu, Minwen Deng 等SC 2024 · 被引用 14 次
- GTuner: tuning DNN computations on GPU via graph attention networkQi Sun, Xinyun Zhang, Hao Geng, Yuxuan Zhao 等DAC 2022 · 被引用 10 次
- Bootstrapping in-situ workflow auto-tuning via combining performance models of component applicationsTong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding 等SC 2021 · 被引用 9 次
- Fast and Efficient DNN Deployment via Deep Gaussian Transfer LearningQi Sun, Chen Bai, Tinghuan Chen, Hao Geng 等ICCV 2021 · 被引用 7 次
相关 Paper
- ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program ReuseJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2022 · 被引用 10 次
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 被引用 27 次
- DeepBurning-SEG: Generating DNN Accelerators of Segment-Grained Pipeline ArchitectureXuyi Cai, Ying Wang, Xiaohan Ma, Yinhe Han 等MICRO 2022 · 被引用 25 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- Understanding and bridging the gaps in current GNN performance optimizationsKezhao Huang, Jidong Zhai, Zhen Zheng, Youngmin Yi 等PPoPP 2021 · 被引用 87 次
