Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
Agathe Balayn, Natasa Rikalo, Jie Yang, Alessandro Bozzon
摘要
Handling failures in computer vision systems that rely on deep learning models remains a challenge. While an increasing number of methods for bug identifcation and correction are proposed, little is known about how practitioners actually search for failures in these models. We perform an empirical study to understand the goals and needs of practitioners, the workfows and artifacts they use, and the challenges and limitations in their process. We interview 18 practitioners by probing them with a carefully crafted failure handling scenario. We observe that there is a great diversity of failure handling workfows in which cooperations are often necessary, that practitioners overlook certain types of failures and bugs, and that they generally do not rely on potentially relevant approaches and tools originally stemming from research. These insights allow to draw a list of research opportunities, such as creating a library of best practices and more representative formalisations of practitioners' goals, developing interfaces to exploit failure handling artifacts, as well as providing specialized training.
• Computing methodologies → Computer vision; • Humancentered computing → Empirical studies in HCI; • Software and its engineering → Software development methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Nurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI-Based SystemsAman Khullar, Nikhil Nalin, Abhishek Prasad, Ann John Mampilli 等CHI 2025 · 被引用 14 次
- Amplifying Rural Educators' Perspectives: A Qualitative Study on the Impacts of Generative AI in Rural U.S. High SchoolsShira Michel, Benjamin Taylor, Sabrina Parra Díaz, Joseph B. Wiggins 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper29
- What is AI Literacy? Competencies and Design ConsiderationsDuri Long, Brian MagerkoCHI 2020 · 被引用 2,947 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 被引用 758 次
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana 等CHI 2020 · 被引用 541 次
- Human Factors in Model Interpretability: Industry Practices, Challenges, and NeedsSungsoo Ray Hong, Jessica Hullman, Enrico BertiniCSCW 2020 · 被引用 219 次
相关 Paper
- How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang 等CHI 2022 · 被引用 22 次
- An empirical study on program failures of deep learning jobsRu Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu 等ICSE 2020 · 被引用 96 次
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein 等CHI 2023 · 被引用 51 次
- Compatibility Issues in Deep Learning Systems: Problems and OpportunitiesJun Wang, Guanping Xiao, Shuai Zhang, Huashan Lei 等FSE 2023 · 被引用 13 次
- A Comprehensive Study of Deep Learning Model Fixing ApproachesHanmo You, Zan Wang, Zishuo Dong, Luanqi Mo 等ICSE 2026
