An Exploratory Study of Deep learning Supply Chain
Xin Tan, Kai Gao, Minghui Zhou, Li Zhang
Abstract
Deep learning becomes the driving force behind many contemporary technologies and has been successfully applied in many fields. Through software dependencies, a multi-layer supply chain (SC) with a deep learning framework as the core and substantial down-stream projects as the periphery has gradually formed and is constantly developing. However, basic knowledge about the structure and characteristics of the SC is lacking, which hinders effective support for its sustainable development. Previous studies on software SC usually focus on the packages in different registries without paying attention to the SCs derived from a single project. We present an empirical study on two deep learning SCs: TensorFlow and PyTorch SCs. By constructing and analyzing their SCs, we aim to understand their structure, application domains, and evolutionary factors. We find that both SCs exhibit a short and sparse hierarchy structure. Overall, the relative growth of new projects increases month by month. Projects have a tendency to attract downstream projects shortly after the release of their packages, later the growth becomes faster and tends to stabilize. We propose three criteria to identify vulnerabilities and identify 51 types of packages and 26 types of projects involved in the two SCs. A comparison reveals their similarities and differences, e.g., TensorFlow SC provides a wealth of packages in experiment result analysis, while PyTorch SC contains more specific framework packages. By fitting the GAM model, we find that the number of dependent packages is significantly negatively associated with the number of downstream projects, but the relationship with the number of authors is nonlinear. Our findings can help further open the "black box" of deep learning SCs and provide insights for their healthy and sustainable development.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software SystemsTrevor Stalnaker, Nathan Wintersgill, Oscar Chaparro, Massimiliano Di Penta et al.ICSE 2024 · 50 citations
- Demystifying Dependency Bugs in Deep Learning StackKaifeng Huang, Bihuan Chen, Susheng Wu, Junming Cao et al.FSE 2023 · 20 citations
- Interoperability in Deep Learning: A User Survey and Failure Analysis of ONNX Model ConvertersPurvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare et al.ISSTA 2024 · 19 citations
- Faster or Slower? Performance Mystery of Python Idioms Unveiled with Empirical EvidenceZejun Zhang, Zhenchang Xing, Xin Xia, Xiwei Xu et al.ICSE 2023 · 16 citations
- PickleBall: Secure Deserialization of Pickle-based Machine Learning ModelsAndreas D. Kellas, Neophytos Christou, Wenxin Jiang, Penghui Li et al.CCS 2025
Related papers
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang et al.FSE 2020 · 121 citations
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- DeepStability: A Study of Unstable Numerical Methods and Their Solutions in Deep LearningEliska Kloberdanz, Kyle G. Kloberdanz, Wei LeICSE 2022 · 16 citations
- Green AI: Do Deep Learning Frameworks Have Different Costs?Stefanos Georgiou, Maria Kechagia, Tushar Sharma, Federica Sarro et al.ICSE 2022 · 90 citations
- Towards Understanding the Faults of JavaScript-Based Deep Learning SystemsLili Quan, Qianyu Guo, Xiaofei Xie, Sen Chen et al.ASE 2022 · 13 citations
