Fast and Efficient DNN Deployment via Deep Gaussian Transfer Learning
Qi Sun, Chen Bai, Tinghuan Chen, Hao Geng, Xinyun Zhang, Yang Bai, Bei Yu
Abstract
Deep neural networks (DNNs) have been widely used recently while their hardware deployment optimizations are very time-consuming and the historical deployment knowledge is not utilized efficiently. In this paper, to accelerate the optimization process and find better deployment configurations, we propose a novel transfer learning method based on deep Gaussian processes (DGPs). Firstly, a deep Gaussian process (DGP) model is built on the historical data to learn empirical knowledge. Secondly, to transfer knowledge to a new task, a tuning set is sampled for the new task under the guidance of the DGP model. Then DGP is tuned according to the tuning set via maximum-a-posteriori (MAP) estimation to accommodate for the new task and finally used to guide the deployments of the task. The experiments show that our method achieves the best inference latencies of convolutions while accelerating the optimization process significantly, compared with previous arts. Preliminaries DNN layers can be represented as several for-loops. Typically, convolutional operations can be represented as a for o in range(0, M): for h in range(0, H): for w in range(0, W): for i in range(0, N): for kh in range(0, KH): for kw in range(0, KW):
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- GTuner: tuning DNN computations on GPU via graph attention networkQi Sun, Xinyun Zhang, Hao Geng, Yuxuan Zhao et al.DAC 2022 · 10 citations
- Glimpse: mathematical embedding of hardware specification for neural compilationByung Hoon Ahn, Sean Kinzer, Hadi EsmaeilzadehDAC 2022 · 4 citations
Builds on9
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark et al.ICLR 2020 · 138 citations
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 102 citations
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 90 citations
Related papers
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsYang Bai, Wenqian Zhao, Shuo Yin, Zixiao Wang et al.EMNLP 2023 · 2 citations
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang et al.DAC 2020 · 15 citations
- Variational Linearized Laplace Approximation for Bayesian Deep LearningLuis A. Ortega Andrés, Simón Rodríguez Santana, Daniel Hernández-LobatoICML 2024 · 12 citations
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson et al.ICML 2023 · 15 citations
- Task-Agnostic Amortized Inference of Gaussian Process HyperparametersSulin Liu, Xingyuan Sun, Peter J. Ramadge, Ryan P. AdamsNeurIPS 2020 · 27 citations
