TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning
Yi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu, Jie Peng, Jianmin Ji, Yanyong Zhang
摘要
Tensor program tuning is a non-convex objective optimization problem, to which search-based approaches have proven to be effective. At the core of the search-based approaches lies the design of the cost model. Though deep learning-based cost models perform significantly better than other methods, they still fall short and suffer from the following problems. First, their feature extraction heavily relies on expert-level domain knowledge in hardware architectures. Even so, the extracted features are often unsatisfactory and require separate considerations for CPUs and GPUs. Second, a cost model trained on one hardware platform usually performs poorly on another, a problem we call cross-hardware unavailability.
In order to address these problems, we propose TLP and MTL-TLP. TLP is a deep learning-based cost model that facilitates tensor program tuning. Instead of extracting features from the tensor program itself, TLP extracts features from the schedule primitives. We treat schedule primitives as tensor languages. TLP is thus a Tensor Language Processing task. In this way, the task of predicting the tensor program latency through the cost model is transformed into a natural language processing (NLP) regression task. MTL-TLP combines Multi-Task Learning and TLP to cope with the crosshardware unavailability problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep LearningYi Zhai, Sijia Yang, Keyu Pan, Renwei Zhang 等OSDI 2024 · 被引用 19 次
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsHanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng 等EuroSys 2024 · 被引用 7 次
- Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program TuningLiang Qiao, Jun Shi, Xiaoyu Hao, Xi Fang 等ASPLOS 2025 · 被引用 5 次
- LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control FlowKaiyan Chang, Wenlong Zhu, Shengwen Liang, Huawei Li 等MICRO 2025 · 被引用 1 次
- FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionZiyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo 等HPCA 2026
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 被引用 90 次
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 被引用 27 次
相关 Paper
- Felix: Optimizing Tensor Programs with Gradient DescentYifan Zhao, Hashim Sharif, Vikram S. Adve, Sasa MisailovicASPLOS 2024 · 被引用 13 次
- Tensor Program Optimization with Probabilistic ProgramsJunru Shao, Xiyou Zhou, Siyuan Feng, Bohan Hou 等NeurIPS 2022 · 被引用 85 次
- Crop: An Analytical Cost Model for Cross-Platform Performance Prediction of Tensor ProgramsXinyu Sun, Yu Zhang, Shuo Liu, Yi ZhaiDAC 2024 · 被引用 2 次
- Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor ProgramsYaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu 等ASPLOS 2023 · 被引用 27 次
- DynaTune: Dynamic Tensor Program Optimization in Deep Neural Network CompilationMinjia Zhang, Menghao Li, Chi Wang, Mingqin LiICLR 2021 · 被引用 18 次
