Interoperability in Deep Learning: A User Survey and Failure Analysis of ONNX Model Converters
Purvish Jajal, Wenxin Jiang, Arav Tewari, Erik Kocinare, Joseph Woo, Anusha Sarraf, Yung-Hsiang Lu, George K. Thiruvathukal, James C. Davis
Abstract
Software engineers develop, fine-tune, and deploy deep learning (DL) models using a variety of development frameworks and runtime environments. DL model converters move models between frameworks and to runtime environments. Conversion errors compromise model quality and disrupt deployment. However, the failure characteristics of DL model converters are unknown, adding risk when using DL interoperability technologies.
This paper analyzes failures in DL model converters. We survey software engineers about DL interoperability tools, use cases, and pain points (N=92). Then, we characterize failures in model converters associated with the main interoperability tool, ONNX (N=200 issues in PyTorch and TensorFlow). Finally, we formulate and test two hypotheses about structural causes for the failures we studied. We find that the node conversion stage of a model converter accounts for ∼75% of the defects, and that 33% of reported failure are related to semantically incorrect models. The cause of semantically incorrect models is elusive, but models with behaviour inconsistencies share operator sequences. Our results motivate future research on making DL interoperability software simpler to maintain, extend, and validate. Research into behavioural tolerances and architectural coverage metrics could be fruitful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d4ac4ff-38e3-4db8-8a04-b089ed5b720eCited by top-tier papers3
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning CompilersSimin Chen, Jinjun Peng, Yixin He, Junfeng Yang et al.S&P 2026 · 11 citations
- Bridging Operator Semantic Inconsistencies: A Source-Level Cross-Framework Model Conversion ApproachXingpei Li, Yan Lei, Zhouyang Jia, Yuanliang Zhang et al.FSE 2025
- PickleBall: Secure Deserialization of Pickle-based Machine Learning ModelsAndreas D. Kellas, Neophytos Christou, Wenxin Jiang, Penghui Li et al.CCS 2025
Builds on11
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu et al.FSE 2020 · 165 citations
- A comprehensive study of autonomous vehicle bugsJoshua Garcia, Yang Feng, Junjie Shen, Sumaya Almanee et al.ICSE 2020 · 127 citations
- A comprehensive study of deep learning compiler bugsQingchao Shen, Haoyang Ma, Junjie Chen, Yongqiang Tian et al.FSE 2021 · 123 citations
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- NNSmith: Generating Diverse and Valid Test Cases for Deep Learning CompilersJiawei Liu, Jinkun Lin, Fabian Ruffy, Cheng Tan et al.ASPLOS 2023 · 90 citations
Related papers
- Improving Deep Learning Framework Testing with Model-Level Metamorphic TestingYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen et al.ISSTA 2025 · 1 citation
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang et al.FSE 2020 · 121 citations
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- DeepStability: A Study of Unstable Numerical Methods and Their Solutions in Deep LearningEliska Kloberdanz, Kyle G. Kloberdanz, Wei LeICSE 2022 · 16 citations
- An empirical study on program failures of deep learning jobsRu Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu et al.ICSE 2020 · 96 citations
