Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance Predictor
Zhen Xie, Murali Emani, Xiaodong Yu, Dingwen Tao, Xin He, Pengfei Su, Keren Zhou, Venkatram Vishwanath
Abstract
For an extended period, graphics processing units (GPUs) have stood as the exclusive choice for training deep neural network (DNN) models. Over time, to serve the growing demands in a more targeted manner, various artificial intelligence-specific hardware, referred to as AI accelerators, have been vigorously developed, aiming to provide more efficient DNN acceleration solutions. However, sufficient solutions are also heterogeneous and thus introduce complexities in accelerator selection. Given a DNN model and a training objective, such as throughput or price-performance ratio, it remains challenging to arrive at the optimal decision among many options due to high reimplementation costs and unexpected performance.
To tackle this challenge, we propose Centimani, a performance predictor that accurately and rapidly predicts DNN training throughput on various AI accelerators, thereby facilitating the accelerator selection process. To achieve this goal, we first analyze typical AI accelerators and draw observations that abstract AI accelerator designs and guide our performance modeling approach. In particular, we construct a memory estimation model and decoupled performance models to select the most appropriate batch size and predict the execution time of DNN training. We validate our approach by applying Centimani to six common DNN models on four typical AI accelerators. Results show that Centimani predicts the throughput with an average accuracy of 93.1% on singledevice training and 90.4% on multiple-device training, thus the optimal accelerator corresponding to the user's training objective can be obtained.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f27ed94-9a7f-4759-9df6-d5b87411f2baCited by top-tier papers1
Ask how each one uses itBuilds on4
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi et al.VLDB 2023 · 113 citations
- Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network TrainingGeoffrey X. Yu, Yubo Gao, Pavel Golikov, Gennady PekhimenkoUSENIX ATC 2021 · 108 citations
- Gradient Compression Supercharged High-Performance Data Parallel DNN TrainingYouhui Bai, Cheng Li, Quan Zhou, Jun Yi et al.SOSP 2021 · 36 citations
- Tahoe: tree structure-aware high performance inference engine for decision tree ensemble on GPUZhen Xie, Wenqian Dong, Jiawen Liu, Hang Liu et al.EuroSys 2021 · 23 citations
Related papers
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 3 citations
- Cambricon-Q: A Hybrid Architecture for Efficient TrainingYongwei Zhao, Chang Liu, Zidong Du, Qi Guo et al.ISCA 2021 · 28 citations
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsHanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng et al.EuroSys 2024 · 7 citations
- Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingYoung H. Oh, Seonghak Kim, Yunho Jin, Sam Son et al.HPCA 2021 · 46 citations
