Understanding Failures of Deep Networks via Robust Feature Extraction
Sahil Singla, Besmira Nushi, Shital Shah, Ece Kamar, Eric Horvitz
Abstract
Traditional evaluation metrics for learned models that report aggregate scores over a test set are insufficient for surfacing important and informative patterns of failure over features and instances. We introduce and study a method aimed at characterizing and explaining failures by identifying visual attributes whose presence or absence results in poor performance. In distinction to previous work that relies upon crowdsourced labels for visual attributes, we leverage the representation of a separate robust model to extract interpretable features and then harness these features to identify failure modes. We further propose a visualization method aimed at enabling humans to understand the meaning encoded in such features and we test the comprehensibility of the features. An evaluation of the methods on the ImageNet dataset demonstrates that: (i) the proposed workflow is effective for discovering important failure modes, (ii) the visualization techniques help humans to understand the extracted features, and (iii) the extracted insights can assist engineers with error analysis and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4a8e201-a9d4-490a-871f-5f61bf150c29Cited by top-tier papers31
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck et al.ICLR 2022 · 178 citations
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Mitigating Spurious Correlations in Multi-modal Models during Fine-tuningYu Yang, Besmira Nushi, Hamid Palangi, Baharan MirzasoleimanICML 2023 · 65 citations
- LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual ImagesViraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy HoffmanNeurIPS 2023 · 59 citations
- A Multimodal Automated Interpretability AgentTamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram et al.ICML 2024 · 57 citations
Builds on2
Related papers
- PRIME: Prioritizing Interpretability in Failure Mode ExtractionKeivan Rezaei, Mehrdad Saberi, Mazda Moayeri, Soheil FeiziICLR 2024 · 9 citations
- Explaining mispredictions of machine learning models using rule inductionJürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali et al.FSE 2021 · 26 citations
- How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang et al.CHI 2022 · 22 citations
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 60 citations
- Distilling Model Failures as Directions in Latent SpaceSaachi Jain, Hannah Lawrence, Ankur Moitra, Aleksander MadryICLR 2023 · 11 citations
