Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
Agathe Balayn, Natasa Rikalo, Jie Yang, Alessandro Bozzon
Abstract
Handling failures in computer vision systems that rely on deep learning models remains a challenge. While an increasing number of methods for bug identifcation and correction are proposed, little is known about how practitioners actually search for failures in these models. We perform an empirical study to understand the goals and needs of practitioners, the workfows and artifacts they use, and the challenges and limitations in their process. We interview 18 practitioners by probing them with a carefully crafted failure handling scenario. We observe that there is a great diversity of failure handling workfows in which cooperations are often necessary, that practitioners overlook certain types of failures and bugs, and that they generally do not rely on potentially relevant approaches and tools originally stemming from research. These insights allow to draw a list of research opportunities, such as creating a library of best practices and more representative formalisations of practitioners' goals, developing interfaces to exploit failure handling artifacts, as well as providing specialized training.
• Computing methodologies → Computer vision; • Humancentered computing → Empirical studies in HCI; • Software and its engineering → Software development methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f4dc041-b90f-4880-9ad9-b08387f8d27eCited by top-tier papers2
- Nurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI-Based SystemsAman Khullar, Nikhil Nalin, Abhishek Prasad, Ann John Mampilli et al.CHI 2025 · 14 citations
- Amplifying Rural Educators' Perspectives: A Qualitative Study on the Impacts of Generative AI in Rural U.S. High SchoolsShira Michel, Benjamin Taylor, Sabrina Parra Díaz, Joseph B. Wiggins et al.CHI 2026 · 2 citations
Builds on29
- What is AI Literacy? Competencies and Design ConsiderationsDuri Long, Brian MagerkoCHI 2020 · 2,947 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- Human Factors in Model Interpretability: Industry Practices, Challenges, and NeedsSungsoo Ray Hong, Jessica Hullman, Enrico BertiniCSCW 2020 · 219 citations
Related papers
- How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang et al.CHI 2022 · 22 citations
- An empirical study on program failures of deep learning jobsRu Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu et al.ICSE 2020 · 96 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
- Compatibility Issues in Deep Learning Systems: Problems and OpportunitiesJun Wang, Guanping Xiao, Shuai Zhang, Huashan Lei et al.FSE 2023 · 13 citations
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo et al.ICSE 2026
