Optimal Kernel Orchestration for Tensor Programs with Korch
Muyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu, Bohan Hou, Gabriele Oliaro, Haojie Wang, Liyan Zheng, Xupeng Miao, Jidong Zhai, Zhihao Jia
摘要
Kernel orchestration is the task of mapping the computation defined in different operators of a deep neural network (DNN) to the execution of GPU kernels on modern hardware platforms. Prior approaches optimize kernel orchestration by greedily applying operator fusion, which fuses the computation of multiple operators into a single kernel, and miss a variety of optimization opportunities in kernel orchestration.
This paper presents Korch, a tensor program optimizer that discovers optimal kernel orchestration strategies for tensor programs. Instead of directly fusing operators, Korch first applies operator fission to decompose tensor operators into a small set of basic tensor algebra primitives. This decomposition enables a diversity of fine-grained, inter-operator optimizations. Next, Korch optimizes kernel orchestration by formalizing it as a constrained optimization problem, leveraging an off-the-shelf binary linear programming solver to discover an optimal orchestration strategy, and generating an executable that can be directly deployed on modern GPU platforms. Evaluation on a variety of DNNs shows that Korch outperforms existing tensor program optimizers by up to 1.7× on V100 GPUs and up to 1.6× on A100 GPUs. Korch is publicly available at https://github.com/humuyan/Korch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Mirage: A Multi-Level Superoptimizer for Tensor ProgramsMengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi 等OSDI 2025 · 被引用 49 次
- MPK: A Compiler and Runtime for Mega-Kernelizing Tensor ProgramsXinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji 等OSDI 2026 · 被引用 20 次
- Elk: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler TechniquesYiqi Liu, Yuqi Xue, Noelle Crawford, Jilong Xue 等MICRO 2025 · 被引用 2 次
- QFactory: Accelerating Quantized Large Language Model Serving with Qtile GraphsQihao Zhang, Mingshu Zhai, Rui Sun, Jidong ZhaiUSENIX ATC 2025 · 被引用 2 次
- VTC: DNN Compilation with Virtual Tensors for Data Movement EliminationMuyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta 等OSDI 2026
它引用的顶会 Paper12
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense PredictionHan Cai, Junyan Li, Muyan Hu, Chuang Gan 等ICCV 2023 · 被引用 265 次
相关 Paper
- EINNET: Optimizing Tensor Programs with Derivation-Based TransformationsLiyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu 等OSDI 2023 · 被引用 19 次
- Felix: Optimizing Tensor Programs with Gradient DescentYifan Zhao, Hashim Sharif, Vikram S. Adve, Sasa MisailovicASPLOS 2024 · 被引用 13 次
- Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor ProgramsYaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu 等ASPLOS 2023 · 被引用 27 次
- ROLLER: Fast and Efficient Tensor Compilation for Deep LearningHongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke 等OSDI 2022 · 被引用 84 次
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng 等PLDI 2021 · 被引用 81 次
