Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of Variance
Hung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier, Jonathan Rosenthal, Lin Tan, Yaoliang Yu, Nachiappan Nagappan
摘要
Deep learning (DL) training algorithms utilize nondeterminism to improve models' accuracy and training efficiency. Hence, multiple identical training runs (e.g., identical training data, algorithm, and network) produce different models with different accuracies and training times. In addition to these algorithmic factors, DL libraries (e.g., TensorFlow and cuDNN) introduce additional variance (referred to as implementation-level variance) due to parallelism, optimization, and floating-point computation. This work is the first to study the variance of DL systems and the awareness of this variance among researchers and practitioners. Our experiments on three datasets with six popular networks show large overall accuracy differences among identical training runs. Even after excluding weak models, the accuracy difference is 10.8%. In addition, implementation-level factors alone cause the accuracy difference across identical training runs to be up to 2.9%, the per-class accuracy difference to be up to 52.4%, and the training time difference to be up to 145.3%. All core libraries (TensorFlow, CNTK, and Theano) and low-level libraries (e.g., cuDNN) exhibit implementation-level variance across all evaluated versions. Our researcher and practitioner survey shows that 83.8% of the 901 participants are unaware of or unsure about any implementation-level variance. In addition, our literature survey shows that only 19.5±3% of papers in recent top software engineering (SE), artificial intelligence (AI), and systems conferences use multiple identical training runs to quantify the variance of their DL approaches. This paper raises awareness of DL variance and directs SE researchers to challenging tasks such as creating deterministic DL implementations to facilitate debugging and improving the reproducibility of DL software and results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li 等ISSTA 2022 · 被引用 142 次
- Deep just-in-time defect prediction: how far are we?Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming ZhangISSTA 2021 · 被引用 97 次
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer 等ICSE 2023 · 被引用 62 次
- Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingShangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu 等NeurIPS 2021 · 被引用 49 次
- Towards Training Reproducible Deep Learning ModelsBoyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin 等ICSE 2022 · 被引用 42 次
它引用的顶会 Paper10
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
- DLFix: context-based code transformation learning for automated program repairYi Li, Shaohua Wang, Tien N. NguyenICSE 2020 · 被引用 201 次
- DeepBillboard: systematic physical-world testing of autonomous driving systemsHusheng Zhou, Wei Li, Zelun Kong, Junfeng Guo 等ICSE 2020 · 被引用 150 次
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 被引用 116 次
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu 等ICSE 2020 · 被引用 101 次
相关 Paper
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu 等FSE 2020 · 被引用 165 次
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 被引用 31 次
- DeepStability: A Study of Unstable Numerical Methods and Their Solutions in Deep LearningEliska Kloberdanz, Kyle G. Kloberdanz, Wei LeICSE 2022 · 被引用 16 次
- Muffin: Testing Deep Learning Libraries via Neural Architecture FuzzingJiazhen Gu, Xuchuan Luo, Yangfan Zhou, Xin WangICSE 2022 · 被引用 63 次
- On the Variance of Neural Network Training with respect to Test Sets and DistributionsKeller JordanICLR 2024 · 被引用 23 次
