UMLAUT: Debugging Deep Learning Programs using Program Structure and Model Behavior
Eldon Schoop, Forrest Huang, Bjoern Hartmann
Abstract
Training deep neural networks can generate non-descriptive error messages or produce unusual output without any explicit errors at all. While experts rely on tacit knowledge to apply debugging strategies, non-experts lack the experience required to interpret model output and correct Deep Learning (DL) programs. In this work, we identify DL debugging heuristics and strategies used by experts, andIn this work, we categorize the types of errors novices run into when writing ML code, and map them onto opportunities where tools could help. We use them to guide the design of Umlaut. Umlaut checks DL program structure and model behavior against these heuristics; provides human-readable error messages to users; and annotates erroneous model output to facilitate error correction. Umlaut links code, model output, and tutorial-driven error messages in a single interface. We evaluated Umlaut in a study with 15 participants to determine its effectiveness in helping developers find and fix errors in their DL programs. Participants using Umlaut found and fixed significantly more bugs and were able to implement fixes for more bugs compared to a baseline condition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6823d323-9867-4282-8c19-b2e00df7ed9dCited by top-tier papers17
- DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsMohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh RajanICSE 2022 · 46 citations
- DeepFD: Automated Fault Diagnosis and Localization for Deep Learning ProgramsJialun Cao, Meiziniu Li, Xiao Chen, Ming Wen et al.ICSE 2022 · 42 citations
- Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual ProgrammingRuofei Du, Na Li, Jing Jin, Michelle Carney et al.CHI 2023 · 33 citations
- Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency AnalysisEldon Schoop, Xin Zhou, Gang Li, Zhourong Chen et al.CHI 2022 · 29 citations
- How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang et al.CHI 2022 · 22 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- Composing Flexibly-Organized Step-by-Step Tutorials from Linked Source Code, Snippets, and OutputsAndrew Head, Jason Jiang, James Smith, Marti A. Hearst et al.CHI 2020 · 29 citations
- Skyline: Interactive In-Editor Computational Performance Profiling for Deep Neural Network TrainingGeoffrey X. Yu, Tovi Grossman, Gennady PekhimenkoUIST 2020 · 16 citations
Related papers
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 62 citations
- A comprehensive study of deep learning compiler bugsQingchao Shen, Haoyang Ma, Junjie Chen, Yongqiang Tian et al.FSE 2021 · 123 citations
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 93 citations
- DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreVincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo TonellaASE 2021 · 41 citations
- Visualizing Examples of Deep Neural Networks at ScaleLitao Yan, Elena L. Glassman, Tianyi ZhangCHI 2021 · 10 citations
