Adaptive Testing of Computer Vision Models
Irena Gao, Gabriel Ilharco, Scott M. Lundberg, Marco Túlio Ribeiro
Abstract
Vision models often fail systematically on groups of data that share common semantic characteristics (e.g., rare objects or unusual scenes), but identifying these failure modes is a challenge. We introduce AdaVision, an interactive process for testing vision models which helps users identify and fix coherent failure modes. Given a natural language description of a coherent group, AdaVision retrieves relevant images from LAION-5B with CLIP. The user then labels a small amount of data for model correctness, which is used in successive retrieval rounds to hill-climb towards high-error regions, refining the group definition. Once a group is saturated, AdaVision uses GPT-3 to suggest new group descriptions for the user to explore. We demonstrate the usefulness and generality of AdaVision in user studies, where users find major bugs in state-of-the-art classification, object detection, and image captioning models. These user-discovered groups have failure rates 2-3x higher than those surfaced by automatic error clustering methods. Finally, finetuning on examples found with AdaVision fixes the discovered bugs when evaluated on unseen examples, without degrading in-distribution accuracy, and while also improving performance on out-of-distribution datasets. * Undertaken in part as an intern at Microsoft Research. collection) and decide if models are safe and fair to deploy [12, 26]. For example, segmentation models for autonomous driving fail in unusual weather. Because we have identified this, we know to deploy such systems with caution and design interventions that simulate diverse weather conditions [39, 49] . Identifying coherent failure modes helps developers make such deployment decisions and design interventions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- Dynamic Evaluation of Large Language Models by Meta Probing AgentsKaijie Zhu, Jindong Wang, Qinlin Zhao, Ruochen Xu et al.ICML 2024 · 65 citations
- Mass-Producing Failures of Multimodal Systems with Language ModelsShengbang Tong, Erik Jones, Jacob SteinhardtNeurIPS 2023 · 54 citations
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 42 citations
- Effective Human-AI Teams via Learned Natural Language Rules and OnboardingHussein Mozannar, Jimin J. Lee, Dennis Wei, Prasanna Sattigeri et al.NeurIPS 2023 · 28 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Discovering Failure Modes of Text-guided Diffusion Models via Adversarial SearchQihao Liu, Adam Kortylewski, Yutong Bai, Song Bai et al.ICLR 2024 · 28 citations
- PRIME: Prioritizing Interpretability in Failure Mode ExtractionKeivan Rezaei, Mehrdad Saberi, Mazda Moayeri, Soheil FeiziICLR 2024 · 9 citations
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 22 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- Error Discovery By Clustering Influence EmbeddingsFulton Wang, Julius Adebayo, Sarah Tan, Diego Garcia-Olano et al.NeurIPS 2023 · 10 citations
