@ CREPE: Can Vision-Language Foundation Models Reason Compositionally?
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, Ranjay Krishna
Abstract
A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, we find that-across 7 architectures trained with 4 algorithms on massive datasets-they struggle at compositionality. To arrive at this conclusion, we introduce a new compositionality evaluation benchmark, CREPE, which measures two important aspects of compositionality identified by cognitive science literature: systematicity and productivity. To measure systematicity, CREPE consists of a test dataset containing over 370K image-text pairs and three different seen-unseen splits. The three splits are designed to test models trained on three popular training datasets: CC-12M, YFCC-15M, and LAION-400M. We also generate 325K, 316K, and 309K hard negative captions for a subset of the pairs. To test productivity, CREPE contains 17K image-text pairs with nine different complexities plus 183K hard negative captions with atomic, swapping and negation foils. The datasets are generated by repurposing the Visual Genome scene graphs and region descriptions and applying handcrafted templates and GPT-3. For systematicity, we find that model performance decreases consistently when novel compositions dominate the retrieval set, with Recall@1 dropping by up to 12%. For productivity, models' retrieval success decays as complexity increases, frequently nearing random chance at high complexity. These results hold regardless of model and training dataset size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09779e1b-2a6d-4315-a12c-8662c3f818d2Cited by top-tier papers73
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang et al.ICCV 2023 · 400 citations
- TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesMaitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim et al.NeurIPS 2024 · 73 citations
- Compositional Chain-of-Thought Prompting for Large Multimodal ModelsChancharik Mitra, Brandon Huang, Trevor Darrell, Roei HerzigCVPR 2024 · 62 citations
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi et al.ICLR 2024 · 53 citations
- Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition LearningWei Li, Hehe Fan, Yongkang Wong, Yi Yang et al.ICML 2024 · 49 citations
Builds on40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky et al.ICLR 2023 · 37 citations
- Text encoders bottleneck compositionality in contrastive vision-language modelsAmita Kamath, Jack Hessel, Kai-Wei ChangEMNLP 2023 · 13 citations
- Interleaved Scene Graphs for Interleaved Text-and-Image Generation AssessmentDongping Chen, Ruoxi Chen, Shu Pu, Zhaoyi Liu et al.ICLR 2025
- SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text GenerationZiyao Xu, Houfeng WangACL 2024
- Iterated Learning Improves Compositionality in Large Vision-Language ModelsChenhao Zheng, Jieyu Zhang, Aniruddha Kembhavi, Ranjay KrishnaCVPR 2024 · 2 citations
