Glimpse: mathematical embedding of hardware specification for neural compilation
Byung Hoon Ahn, Sean Kinzer, Hadi Esmaeilzadeh
Abstract
Success of Deep Neural Networks (DNNs) and their computational intensity has heralded Cambrian explosion of DNN hardware. While hardware design has advanced significantly, optimizing the code for them is still an open challenge. Recent research has moved past traditional compilation techniques and taken a stochastic search algorithmic path that blindly generates rather stochastic samples of the binaries for real hardware measurements to guide the search. This paper opens a new dimension by incorporating the mathematical embedding of the hardware specification of the GPU accelerators dubbed Blueprint to better guide the search algorithm and focus on sub-spaces that have higher potential for yielding higher performance binaries. While various sample efficient yet blind hardware-agnostic techniques have been proposed, none of the state-of-the-art compilers have considered hardware specification as hints to improve the sample efficiency and the search. To mathematically embed the hardware specifications into the search, we devise a Bayesian optimization framework called Glimpse with multiple exclusively unique components. We first use the Blueprint as an input to generate prior distributions of different dimensions in the search space. Then, we devise a light-weight neural acquisition function that takes into account the Blueprint to conform to the hardware specification while balancing the exploration-exploitation trade-off. Finally, we generate an ensemble of predictors from the Blueprint that collectively vote to reject invalid binary samples. We compare Glimpse with hardware-agnostic compilers. Comparison to AutoTVM [3], Chameleon [2], and DGP [16] with multiple generations of GPUs shows that Glimpse provides 6.73×, 1.51×, and 1.92× faster compilation time, respectively, while also achieving the best inference latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d11cc59d-e40a-4f1c-bc13-cc1213ef3639Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr et al.ICLR 2020 · 104 citations
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 90 citations
- AdaTune: Adaptive Tensor Program Compilation Made EfficientMenghao Li, Minjia Zhang, Chi Wang, Mingqin LiNeurIPS 2020 · 39 citations
- DynaTune: Dynamic Tensor Program Optimization in Deep Neural Network CompilationMinjia Zhang, Menghao Li, Chi Wang, Mingqin LiICLR 2021 · 18 citations
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang et al.DAC 2020 · 15 citations
Related papers
- Bayesian Code Diffusion for Efficient Automatic Deep Learning Program OptimizationIsu Jeong, Seulki LeeOSDI 2025
- Leveraging Domain Information for the Efficient Automated Design of Deep Learning AcceleratorsChirag Sakhuja, Zhan Shi, Calvin LinHPCA 2023 · 9 citations
- Romou: rapidly generate high-performance tensor kernels for mobile GPUsRendong Liang, Ting Cao, Jicheng Wen, Manni Wang et al.MobiCom 2022 · 13 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- EDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI SolutionsYuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu et al.DAC 2020 · 79 citations
