Mass-Producing Failures of Multimodal Systems with Language Models
Shengbang Tong, Erik Jones, Jacob Steinhardt
Abstract
Deployed multimodal systems can fail in ways that evaluators did not anticipate. In order to find these failures before deployment, we introduce MULTIMON, a system that automatically identifies systematic failures-generalizable, naturallanguage descriptions of patterns of model failures. To uncover systematic failures, MULTIMON scrapes a corpus for examples of erroneous agreement: inputs that produce the same output, but should not. It then prompts a language model (e.g., GPT-4) to find systematic patterns of failure and describe them in natural language. We use MULTIMON to find 14 systematic failures (e.g., "ignores quantifiers") of the CLIP text-encoder, each comprising hundreds of distinct inputs (e.g., "a shelf with a few/many books"). Because CLIP is the backbone for most state-of-the-art multimodal models, these inputs produce failures in Midjourney 5.1, DALL-E, VideoFusion, and others. MULTIMON can also steer towards failures relevant to specific use cases, such as self-driving cars. We see MULTIMON as a step towards evaluation that autonomously explores the long tail of potential system failures. 2 * Equal contribution 2 Code for MULTIMON is available at https://github.com/tsb0601/MultiMon 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13bc9efa-4328-4afb-84d3-88afd0586ff7Cited by top-tier papers23
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma et al.CVPR 2024 · 111 citations
- Concept Bottleneck Generative ModelsAya Abdelsalam Ismail, Julius Adebayo, Héctor Corrada Bravo, Stephen Ra et al.ICLR 2024 · 43 citations
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- Understanding and Improving Training-free Loss-based Diffusion GuidanceYifei Shen, Xinyang Jiang, Yifan Yang, Yezhen Wang et al.NeurIPS 2024 · 36 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Discovering Failure Modes of Text-guided Diffusion Models via Adversarial SearchQihao Liu, Adam Kortylewski, Yutong Bai, Song Bai et al.ICLR 2024 · 28 citations
- Symbal: Detecting Systematic Misalignments in Model-Generated CaptionsMaya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari et al.ICML 2026
- Text encoders bottleneck compositionality in contrastive vision-language modelsAmita Kamath, Jack Hessel, Kai-Wei ChangEMNLP 2023 · 13 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & TasksLetitia Parcalabescu, Anette FrankACL 2023 · 15 citations
