COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real Images
Ben Bogin, Shivanshu Gupta, Matt Gardner, Jonathan Berant
Abstract
While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images. In this work, we propose COVR, a new test-bed for visually-grounded compositional generalization with real images. To create COVR, we use real images annotated with scene graphs, and propose an almost fully automatic procedure for generating question-answer pairs along with a set of context images. COVR focuses on questions that require complex reasoning, including higherorder operations such as quantification and aggregation. Due to the automatic generation process, COVR facilitates the creation of compositional splits, where models at test time need to generalize to new concepts and compositions in a zero-or few-shot setting. We construct compositional splits using COVR and demonstrate a myriad of cases where state-ofthe-art pre-trained language-and-vision models struggle to compositionally generalize.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9cafa72-ff8f-44f8-b70e-b6d68ab1c150Cited by top-tier papers7
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsSreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi et al.ICLR 2024 · 53 citations
- Image Retrieval from Contextual DescriptionsBenno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal et al.ACL 2022 · 37 citations
- When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky et al.ICLR 2023 · 37 citations
- When are Lemons Purple? The Concept Association Bias of Vision-Language ModelsYingtian Tang, Yutaro Yamada, Yoyo Zhang, Ilker YildirimEMNLP 2023 · 9 citations
Builds on6
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionZhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong et al.CVPR 2020
Related papers
- Generative Compositional Augmentations for Scene Graph PredictionBoris Knyazev, Harm de Vries, Catalina Cangea, Graham W. Taylor et al.ICCV 2021 · 30 citations
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie et al.CVPR 2024 · 4 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris et al.CVPR 2021
- V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive MatricesDamien Teney, Peng Wang, Jiewei Cao, Lingqiao Liu et al.AAAI 2020 · 37 citations
