Towards Training Reproducible Deep Learning Models
Boyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin, Gopi Krishnan Rajbahadur, Zhen Ming Jiang
Abstract
Reproducibility is an increasing concern in Artificial Intelligence (AI), particularly in the area of Deep Learning (DL). Being able to reproduce DL models is crucial for AI-based systems, as it is closely tied to various tasks like training, testing, debugging, and auditing. However, DL models are challenging to be reproduced due to issues like randomness in the software (e.g., DL algorithms) and non-determinism in the hardware (e.g., GPU). There are various practices to mitigate some of the aforementioned issues. However, many of them are either too intrusive or can only work for a specific usage context. In this paper, we propose a systematic approach to training reproducible DL models. Our approach includes three main parts: (1) a set of general criteria to thoroughly evaluate the reproducibility of DL models for two different domains, (2) a unified framework which leverages a record-and-replay technique to mitigate software-related randomness and a profile-and-patch technique to control hardware-related non-determinism, and (3) a reproducibility guideline which explains the rationales and the mitigation strategies on conducting a reproducible training process for DL models. Case study results show our approach can successfully reproduce six open source and one commercial DL models. CCS CONCEPTS • Software and its engineering → Empirical software validation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f21cb20-beef-43ad-a460-0851447be660Cited by top-tier papers6
- FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingZiqi Zhang, Yuanchun Li, Bingyan Liu, Yifeng Cai et al.ICSE 2023 · 8 citations
- Revealing Floating-Point Accumulation Orders in Software/Hardware ImplementationsPeichen Xie, Yanjie Gao, Yang Wang, Jilong XueUSENIX ATC 2025 · 6 citations
- Reproduce, Replicate, Reevaluate. The Long but Safe Way to Extend Machine Learning MethodsLuisa Werner, Nabil Layaïda, Pierre Genevès, Jérôme Euzenat et al.AAAI 2024 · 2 citations
- Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered AgentsBenjamin Rombaut, Sogol Masoumzadeh, Kirill Vasilevski, Dayi Lin et al.ASE 2025 · 1 citation
- Synchronizing Probabilities in Model-Driven Lossless CompressionAviv Adler, Jennifer TangICLR 2026 · 1 citation
Builds on2
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 65 citations
Related papers
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent AgentMehil Shah, Mohammad Masudur Rahman, Foutse KhomhICSE 2026
- Evaluating Deep Neural Networks in Deployment: A Comparative Study (Replicability Study)Eduard Pinconschi, Divya Gopinath, Rui Abreu, Corina S. PasareanuISSTA 2024
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo et al.ICSE 2026
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer et al.ICSE 2023 · 62 citations
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu et al.FSE 2020 · 165 citations
