Ansor: Generating High-Performance Tensor Programs for Deep Learning
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica
Abstract
High-performance tensor programs are crucial to guarantee efficient execution of deep learning models. However, obtaining performant tensor programs for different operators on various hardware platforms is notoriously difficult. Currently, deep learning systems rely on vendor-provided kernel libraries or various search strategies to get performant tensor programs. These approaches either require significant engineering efforts in developing platform-specific optimization code or fall short in finding high-performance programs due to restricted search space and ineffective exploration strategy. We present Ansor, a tensor program generation framework for deep learning applications. Compared with existing search strategies, Ansor explores much more optimization combinations by sampling programs from a hierarchical representation of the search space. Ansor then fine-tunes the sampled programs with evolutionary search and a learned cost model to identify the best programs. Ansor can find high-performance programs that are outside the search space of existing state-of-the-art approaches. Besides, Ansor utilizes a scheduler to simultaneously optimize multiple subgraphs in a set of deep neural networks. Our evaluation shows that Ansor improves the execution performance of deep neural networks on the Intel CPU, ARM CPU, and NVIDIA GPU by up to , , and , respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ad7b489-8de4-4cb0-90e7-d341b6feda61Cited by top-tier papers132
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 220 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationCong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng et al.ISCA 2023 · 151 citations
- Lite Pose: Efficient Architecture Design for 2D Human Pose EstimationYihan Wang, Muyang Li, Han Cai, Wei-Ming Chen et al.CVPR 2022 · 117 citations
Builds on2
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- FeatGraph: a flexible and efficient backend for graph neural network systemsYuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu et al.SC 2020 · 57 citations
Related papers
- Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep LearningYi Zhai, Sijia Yang, Keyu Pan, Renwei Zhang et al.OSDI 2024 · 19 citations
- Felix: Optimizing Tensor Programs with Gradient DescentYifan Zhao, Hashim Sharif, Vikram S. Adve, Sasa MisailovicASPLOS 2024 · 13 citations
- Optimal Kernel Orchestration for Tensor Programs with KorchMuyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu et al.ASPLOS 2024 · 11 citations
- Bayesian Code Diffusion for Efficient Automatic Deep Learning Program OptimizationIsu Jeong, Seulki LeeOSDI 2025
- EINNET: Optimizing Tensor Programs with Derivation-Based TransformationsLiyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu et al.OSDI 2023 · 19 citations
