Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor Programs
Yaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu, Yida Wang, Gennady Pekhimenko
摘要
As deep learning models nowadays are widely adopted by both cloud services and edge devices, reducing the latency of deep learning model inferences becomes crucial to provide efficient model serving. However, it is challenging to develop efficient tensor programs for deep learning operators due to the high complexity of modern accelerators (e.g., NVIDIA GPUs and Google TPUs) and the rapidly growing number of operators.
Deep learning compilers, such as Apache TVM, adopt declarative scheduling primitives to lower the bar of developing tensor programs. However, we show that this approach is insufficient to cover state-of-the-art tensor program optimizations (e.g., double buffering). In this paper, we propose to embed the scheduling process into tensor programs and use dedicated mappings, called task mappings, to define the computation assignment and ordering directly in the tensor programs. This new approach greatly enriches the expressible optimizations by allowing developers to manipulate tensor programs at a much finer granularity (e.g., allowing programstatement-level optimizations). We call the proposed method the task-mapping programming paradigm. In addition, we propose a new post-scheduling fusion optimization that allows developers to focus on scheduling every single operator and automates the fusion after scheduling. It greatly reduces the engineering efforts for operator fusion. Our proposed paradigm also constructs an efficient hardware-centric schedule space, which is agnostic to the program input size and greatly reduces the tuning time.
With the proposed paradigm, we implement a deep learning compiler -Hidet. Extensive experiments on modern convolution and transformer models show that Hidet outperforms state-of-theart DNN inference framework, ONNX Runtime, and compiler, TVM equipped with scheduler AutoTVM and Ansor, by up to 1.48× (1.22× * Part of the work done while interning at Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- EINNET: Optimizing Tensor Programs with Derivation-Based TransformationsLiyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu 等OSDI 2023 · 被引用 19 次
- Optimal Kernel Orchestration for Tensor Programs with KorchMuyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu 等ASPLOS 2024 · 被引用 11 次
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 9 次
- Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-OptimizationZhanda Zhu, Christina Giannoula, Muralidhar Andoorveedu, Qidong Su 等EuroSys 2025 · 被引用 8 次
它引用的顶会 Paper15
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- Dynamic Tensor RematerializationMarisa Kirisame, Steven Lyubomirsky, Altan Haan, Jennifer Brennan 等ICLR 2021 · 被引用 115 次
相关 Paper
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsYang Bai, Wenqian Zhao, Shuo Yin, Zixiao Wang 等EMNLP 2023 · 被引用 2 次
- TLP: A Deep Learning-Based Cost Model for Tensor Program TuningYi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu 等ASPLOS 2023 · 被引用 42 次
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsHanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng 等EuroSys 2024 · 被引用 7 次
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang 等ASPLOS 2024 · 被引用 14 次
- Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloadsEvangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman 等SC 2021 · 被引用 2 次
