ETS: Deep Learning Training Iteration Time Prediction based on Execution Trace Sliding Window
Zichao Yang, Hao Guo, Heng Wu, Yuewen Wu, Hua Zhong, Wenbo Zhang, Chuan Zhou, Yan Liu
Abstract
Deep learning (DL) has become essential across various computer science domains. Accurately predicting iteration time for DL models in diverse cloud data center environments is critical for making high-quality scheduling decisions. Existing approaches neglect the sequential features inherent in the runtime execution, leading to issues such as overlooking DL framework overhead and struggling to handle diverse sizes of DL models, resulting in either low accuracy or slow convergence of the prediction model. This paper introduces ETS, a novel iteration time prediction method utilizing execution trace sliding windows. Our observation reveals that DL models exhibit a highly sequential runtime execution nature. Building upon this insight, we leverage sliding windows to extract a novel type of sequential features from the runtime execution trace. These features comprehensively capture DL framework overhead and address the diversity challenge in DL model sizes. By combining a best-practice method to train a prediction model, we achieve high accuracy and rapid convergence simultaneously. Experimental validation on over 14,000 DL model configurations demonstrates ETS's effectiveness in predicting the iteration time of DL models, achieving a mere 5.9% prediction error with a training time at the 10-minute level, and improving scheduling outcomes by reducing job completion time by 17%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU ClustersZiyue Luo, Jia Liu, Myungjin Lee, Ness B. ShroffINFOCOM 2025 · 5 citations
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model ExplorationHairui Zhao, Hongliang Li, Qi Tian, Jie Wu et al.INFOCOM 2025 · 3 citations
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams et al.SC 2020 · 21 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
