A Comprehensive Study of Real-World Bugs in Machine Learning Model Optimization
Hao Guan, Ying Xiao, Jiaying Li, Yepang Liu, Guangdong Bai
摘要
Due to the great advance in machine learning (ML) techniques, numerous ML models are expanding their application domains in recent years. To adapt for resource-constrained platforms such as mobile and Internet of Things (IoT) devices, pre-trained models are often processed to enhance their efficiency and compactness, using optimization techniques such as pruning and quantization. Similar to the optimization process in other complex systems, e.g., program compilers and databases, optimizations for ML models can contain bugs, leading to severe consequences such as system crashes and financial loss. While bugs in training, compiling and deployment stages have been extensively studied, there is still a lack of systematic understanding and characterization of model optimization bugs (MOBs). In this work, we conduct the first empirical study to identify and characterize MOBs. We collect a comprehensive dataset containing 371 MOBs from TensorFlow and PyTorch, the most extensively used open-source ML frameworks, covering the entire development time span of their optimizers (May 2019 to August 2022). We then investigate the collected bugs from various perspectives, including their symptoms, root causes, life cycles, detection and fixes. Our work unveils the status quo of MOBs in the wild, and reveals their features on which future detection techniques can be based. Our findings also serve as a warning to the developers and the users of ML frameworks, and an appeal to our research community to enact dedicated countermeasures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Compatibility Issues in Deep Learning Systems: Problems and OpportunitiesJun Wang, Guanping Xiao, Shuai Zhang, Huashan Lei 等FSE 2023 · 被引用 13 次
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning CompilersSimin Chen, Jinjun Peng, Yixin He, Junfeng Yang 等S&P 2026 · 被引用 11 次
- Understanding Transaction Bugs in Database SystemsZiyu Cui, Wensheng Dou, Yu Gao, Dong Wang 等ICSE 2024 · 被引用 9 次
- Large Language Models Can Connect the Dots: Exploring Model Optimization Bugs with Domain Knowledge-Aware PromptsHao Guan, Guangdong Bai, Yepang LiuISSTA 2024 · 被引用 8 次
- PolyJuice: Detecting Mis-compilation Bugs in Tensor Compilers with Equality Saturation Based RewritingChijin Zhou, Bingzhou Qian, Gwihwan Go, Quan Zhang 等OOPSLA 2024 · 被引用 7 次
它引用的顶会 Paper11
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio 等ICSE 2020 · 被引用 281 次
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu 等FSE 2020 · 被引用 165 次
- A comprehensive study of deep learning compiler bugsQingchao Shen, Haoyang Ma, Junjie Chen, Yongqiang Tian 等FSE 2021 · 被引用 123 次
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang 等FSE 2020 · 被引用 121 次
- Detecting optimization bugs in database engines via non-optimizing reference engine constructionManuel Rigger, Zhendong SuFSE 2020 · 被引用 104 次
相关 Paper
- Understanding performance problems in deep learning systemsJunming Cao, Bihuan Chen, Chao Sun, Longjie Hu 等FSE 2022 · 被引用 33 次
- Detecting TensorFlow Program Bugs in Real-World Industrial EnvironmentChen Liu, Jie Lu, Guangwei Li, Ting Yuan 等ASE 2021 · 被引用 13 次
- Understanding the Bug Characteristics and Fix Strategies of Federated Learning SystemsXiaohu Du, Xiao Chen, Jialun Cao, Ming Wen 等FSE 2023 · 被引用 6 次
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile ApplicationsZhenpeng Chen, Huihan Yao, Yiling Lou, Yanbin Cao 等ICSE 2021 · 被引用 73 次
- Improving Deep Learning Framework Testing with Model-Level Metamorphic TestingYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen 等ISSTA 2025 · 被引用 1 次
