Image Retrieval from Contextual Descriptions
Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Maria Ponti, Siva Reddy
Abstract
The ability to integrate context, including perceptual and temporal cues, plays a pivotal role in grounding the meaning of a linguistic utterance. In order to measure to what extent current vision-and-language models master this ability, we propose a new multimodal challenge, Image Retrieval from Contextual Descriptions (IMAGECODE). In particular, models are tasked with retrieving the correct image from a set of 10 minimally contrastive candidates based on a contextual description. As such, each description contains only the details that help distinguish between images. Because of this, descriptions tend to be complex in terms of syntax and discourse and require drawing pragmatic inferences. Images are sourced from both static pictures and video frames. We benchmark several state-of-the-art models, including both cross-encoders such as ViLBERT and bi-encoders such as CLIP, on IMAGECODE. Our results reveal that these models dramatically lag behind human performance: the best variant achieves an accuracy of 20.9 on video frames and 59.4 on static pictures, compared with 90.8 in humans. Furthermore, we experiment with new model variants that are better equipped to incorporate visual and temporal context into their representations, which achieve modest gains. Our hope is that IMAGECODE will foster progress in grounded language understanding by encouraging models to focus on fine-grained visual differences. We make code and dataset publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementXiangyang Zhu, Renrui Zhang, Bowei He, Aojun Zhou et al.ICCV 2023 · 121 citations
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras et al.NeurIPS 2025 · 97 citations
- Equivariant Similarity for Vision-Language Foundation ModelsTan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin et al.ICCV 2023 · 67 citations
- Improving fine-grained understanding in image-text pre-trainingIoana Bica, Anastasija Ilic, Matthias Bauer, Goker Erdogan et al.ICML 2024 · 53 citations
- VisMin: Visual Minimal-Change UnderstandingRabiul Awal, Saba Ahmadi, Le Zhang, Aishwarya AgrawalNeurIPS 2024 · 25 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy et al.EMNLP 2021 · 87 citations
- COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real ImagesBen Bogin, Shivanshu Gupta, Matt Gardner, Jonathan BerantEMNLP 2021 · 13 citations
- Image Change Captioning by Learning From an Auxiliary TaskMehrdad Hosseinzadeh, Yang WangCVPR 2021
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic et al.CVPR 2021
Related papers
- CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang et al.ACL 2024 · 3 citations
- Accommodating Audio Modality in CLIP for Multimodal ProcessingLudan Ruan, Anwen Hu, Yuqing Song, Liang Zhang et al.AAAI 2023 · 18 citations
- VISTA: Visualized Text Embedding For Universal Multi-Modal RetrievalJunjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao et al.ACL 2024
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao et al.AAAI 2026 · 2 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
