Generative Zero-Shot Composed Image Retrieval
Lan Wang, Wei Ao, Vishnu Naresh Boddeti, Ser-Nam Lim
Abstract
Composed Image Retrieval (CIR) is a vision-language task utilizing queries comprising images and textual descriptions to achieve precise image retrieval. This task seeks to find images that are visually similar to a reference image while incorporating specific changes or features described textually (visual delta). CIR enables a more flexible and user-specific retrieval by bridging visual data with verbal instructions. This paper introduces a novel generative method that augments Composed Image Retrieval by Composed Image Generation (CIG) to provide pseudotarget images. CIG utilizes a textual inversion network to map reference images into semantic word space, which generates pseudo-target images in combination with textual descriptions. These images serve as additional visual information, significantly improving the accuracy and relevance of retrieved images when integrated into existing retrieval frameworks. Experiments conducted across multiple CIR datasets and several baseline methods demonstrate improvements in retrieval performance, which shows the potential of our approach as an effective add-on for existing composed image retrieval. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84498085-c582-4dbd-b57e-9b48de0068c3Cited by top-tier papers4
- Conditional Representation Learning for Customized TasksHonglin Liu, Chao Sun, Peng Hu, Yunfan Li et al.NeurIPS 2025 · 6 citations
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image RetrievalTianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao et al.CVPR 2026 · 4 citations
- Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningHaomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng et al.AAAI 2026 · 1 citation
- Modality and Task Adaptation for Enhanced Zero-shot Composed Image RetrievalHaiwen Li, Delong Liu, Zhaohui Hou, Zeliang Ma et al.AAAI 2026 · 1 citation
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Visual Delta Generator with Large Multi-Modal Models for Semi-Supervised Composed Image RetrievalYoung Kyun Jang, Donghyun Kim, Zihang Meng, Dat Huynh et al.CVPR 2024
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu et al.SIGIR 2024 · 29 citations
- Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware PerspectiveZhe Li, Lei Zhang, Kun Zhang, Weidong Chen et al.SIGIR 2025 · 5 citations
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li et al.CVPR 2023
