Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
You Li, Fan Ma, Yi Yang
Abstract
The Zero-shot Composed Image Retrieval (ZSCIR) requires retrieving images that match the query image and the relative captions. Current methods focus on projecting the query image into the text feature space, subsequently combining them with features of query texts for retrieval. However, retrieving images only with the text features cannot guarantee detailed alignment due to the natural gap between images and text. In this paper, we introduce Imagined Proxy for CIR (IP-CIR), a training-free method that creates a proxy image aligned with the query image and text description, enhancing query representation in the retrieval process. We first leverage the large language model's generalization capability to generate an image layout, and then apply both the query text and image for conditional generation. The robust query features are enhanced by merging the proxy image, query image, and text semantic perturbation. Our newly proposed balancing metric integrates text-based and proxy retrieval similarities, allowing for more accurate retrieval of the target image while incorporating image-side information into the process. Experiments on three public datasets demonstrate that our method significantly improves retrieval performances. We achieve state-of-the-art (SOTA) results on the CIRR dataset with a Recall@K of 70.07 at K=10. Additionally, we achieved an improvement in Recall@10 on the FashionIQ dataset, rising from 45.11 to 45.74, and improved the baseline performance in CIRCO with a mAPK@10 score, increasing from 32.24 to 34.26.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image ModelsDewei Zhou, Mingwei Li, Zongxin Yang, Yi YangICCV 2025 · 5 citations
- GenIR: Generative Visual Feedback for Mental Image RetrievalDiji Yang, Minghao Liu, Chung-Hsiang Lo, Yi Zhang et al.NeurIPS 2025 · 4 citations
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image RetrievalTianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao et al.CVPR 2026 · 4 citations
- STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image RetrievalMiaoge Li, Dongsheng Wang, Zening Sun, Jinsen Zhang et al.CVPR 2026 · 2 citations
- FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured ScriptsYou Li, Dewei Zhou, Fan Ma, Fu Li et al.CVPR 2026 · 2 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Semantic Editing Increment Benefits Zero-Shot Composed Image RetrievalZhenyu Yang, Shengsheng Qian, Dizhan Xue, Jiahong Wu et al.ACM MM 2024 · 15 citations
- LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image RetrievalZhenyu Yang, Dizhan Xue, Shengsheng Qian, Weiming Dong et al.SIGIR 2024 · 52 citations
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 214 citations
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li et al.CVPR 2023
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
