An empirical study on program failures of deep learning jobs
Ru Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu, Haoxiang Lin, Mao Yang
Abstract
Deep learning has made significant achievements in many application areas. To train and test models more efficiently, enterprise developers submit and run their deep learning programs on a shared, multi-tenant platform. However, some of the programs fail after a long execution time due to code/script defects, which reduces the development productivity and wastes expensive resources such as GPU, storage, and network I/O. This paper presents the first comprehensive empirical study on program failures of deep learning jobs. 4960 real failures are collected from a deep learning platform in Microsoft. We manually examine their failure messages and classify them into 20 categories. In addition, we identify the common root causes and bug-fix solutions on a sample of 400 failures. To better understand the current testing and debugging practices for deep learning, we also conduct developer interviews. Our major findings include: (1) 48.0% of the failures occur in the interaction with the platform rather than in the execution of code logic, mostly due to the discrepancies between local and platform execution environments; (2) Deep learning specific failures (13.5%) are mainly caused by inappropriate model parameters/structures and framework API misunderstanding; (3) Current debugging practices are not efficient for fault localization in many cases, and developers need more deep learning specific tools. Based on our findings, we further suggest possible research topics and tooling support that could facilitate future deep learning development. CCS CONCEPTS • Software and its engineering → Software defect analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2592dfd3-901c-4db0-9194-52b4411101a8Cited by top-tier papers35
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- IoT Bugs and Development ChallengesAmir Makhshari, Ali MesbahICSE 2021 · 76 citations
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile ApplicationsZhenpeng Chen, Huihan Yao, Yiling Lou, Yanbin Cao et al.ICSE 2021 · 73 citations
Related papers
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao et al.ICSE 2024 · 22 citations
- Understanding performance problems in deep learning systemsJunming Cao, Bihuan Chen, Chao Sun, Longjie Hu et al.FSE 2022 · 33 citations
- Compatibility Issues in Deep Learning Systems: Problems and OpportunitiesJun Wang, Guanping Xiao, Shuai Zhang, Huashan Lei et al.FSE 2023 · 13 citations
- Towards Understanding the Faults of JavaScript-Based Deep Learning SystemsLili Quan, Qianyu Guo, Xiaofei Xie, Sen Chen et al.ASE 2022 · 13 citations
- Detecting TensorFlow Program Bugs in Real-World Industrial EnvironmentChen Liu, Jie Lu, Guangwei Li, Ting Yuan et al.ASE 2021 · 13 citations
