Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, Candace Ross
Abstract
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly-but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set of fine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab0aaa72-ed2a-4c8d-8f5b-7e14e6bcca71Cited by top-tier papers190
- Your Diffusion Model is Secretly a Zero-Shot ClassifierAlexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown et al.ICCV 2023 · 341 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- What You See is What You Read? Improving Text-Image Alignment EvaluationMichal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni et al.NeurIPS 2023 · 147 citations
- Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model EvaluationShih-Ying Yeh, Yu-Guan Hsieh, Zhidong Gao, Bernard B. W. Yang et al.ICLR 2024 · 133 citations
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 120 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
Related papers
- Why is Winoground Hard? Investigating Failures in Visuolinguistic CompositionalityAnuj Diwan, Layne Berry, Eunsol Choi, David Harwath et al.EMNLP 2022 · 15 citations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie et al.CVPR 2024 · 4 citations
- The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsMostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov et al.ACL 2020 · 1 citation
- Going Beyond Nouns With Vision & Language Models Using Synthetic DataPaola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh et al.ICCV 2023 · 49 citations
