CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor Programs
Hanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng, Yibo Zhu, Haibin Lin, Chuan Wu
Abstract
Deep Neural Networks (DNNs) have shown excellent performance in a wide range of machine learning applications. Knowing the latency of running a DNN model or tensor program on a specific device is useful in various tasks, such as DNN graph- or tensor-level optimization and device selection. Considering the large space of DNN models and devices that impedes direct profiling of all combinations, recent efforts focus on building a predictor to model the performance of DNN models on different devices. However, none of the existing attempts have achieved a cost model that can accurately predict the performance of various tensor programs while supporting both training and inference accelerators. We propose CDMPP, an efficient tensor program latency prediction framework for both cross-model and cross-device prediction. We design an informative but efficient representation of tensor programs, called compact ASTs, and a pre-order-based positional encoding method, to capture the internal structure of tensor programs. We develop a domain-adaption-inspired method to learn domain-invariant representations and devise a KMeans-based sampling algorithm, for the predictor to learn from different domains (i.e., different DNN operators and devices). Our extensive experiments on a diverse range of DNN models and devices demonstrate that CDMPP significantly outperforms state-of-the-art baselines with 14.03% and 10.85% prediction error for cross-model and cross-device prediction, respectively, and one order of magnitude higher training efficiency. The implementation and the expanded dataset are available at https://github.com/joapolarbear/cdmpp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d6f95c7-ba91-4583-af63-884a35689a73Builds on11
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 672 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- BRP-NAS: Prediction-based NAS using GCNsLukasz Dudziak, Thomas Chau, Mohamed S. Abdelfattah, Royson Lee et al.NeurIPS 2020 · 233 citations
Related papers
- TLP: A Deep Learning-Based Cost Model for Tensor Program TuningYi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu et al.ASPLOS 2023 · 42 citations
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsYang Bai, Wenqian Zhao, Shuo Yin, Zixiao Wang et al.EMNLP 2023 · 2 citations
- Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor ProgramsYaoyao Ding, Cody Hao Yu, Bojian Zheng, Yizhi Liu et al.ASPLOS 2023 · 27 citations
- Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule SearchZhiqi Lin, Youshan Miao, Guanbin Xu, Cheng Li et al.HPCA 2024 · 6 citations
- Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloadsEvangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman et al.SC 2021 · 2 citations
