Language-only Efficient Training of Zero-shot Composed Image Retrieval
Geonmo Gu, Sanghyuk Chun, Wonjae Kim, Yoohoon Kang, Sangdoo Yun
Abstract
Composed image retrieval (CIR) task takes a composed query of image and text, aiming to search relative images for both conditions. Conventional CIR approaches need a training dataset composed of triplets of query image, query text, and target image, which is very expensive to collect. Several recent works have worked on the zero-shot (ZS) CIR paradigm to tackle the issue without using pre-collected triplets. However, the existing ZS-CIR methods show limited backbone scalability and generalizability due to the lack of diversity of the input texts during training. We propose a novel CIR framework, only using language for its training. Our LinCIR (Language-only training for CIR) can be trained only with text datasets by a novel self-supervision named self-masking projection (SMP). We project the text latent embedding to the token embedding space and construct a new text by replacing the keyword tokens of the original text. Then, we let the new and original texts have the same latent embedding vector. With this simple strategy, LinCIR is surprisingly efficient and highly effective; LinCIR with CLIP ViT-G backbone is trained in 48 minutes and shows the best ZS-CIR performances on four different CIR benchmarks, CIRCO, GeneCIS, FashionIQ, and CIRR, even outperforming supervised method on FashionIQ. Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5847e96f-252a-4166-b454-4aa36e196e26Cited by top-tier papers13
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 26 citations
- Towards Cross-Modal Backward-Compatible Representation Learning for Vision-Language ModelsYoung Kyun Jang, Ser-Nam LimICCV 2025 · 3 citations
- An Efficient Post-Hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image RetrievalJaeseok Byun, Seokhyeon Jeong, Wonjae Kim, Sanghyuk Chun et al.ICCV 2025 · 2 citations
- G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image RetrievalJiyoung Lim, Heejae Yang, Jee-Hyong LeeCVPR 2026 · 1 citation
- XR: Cross-Modal Agents for Composed Image RetrievalZhongyu Yang, Wei Pang, Yingfang YuanWWW 2026 · 1 citation
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
Related papers
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 214 citations
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li et al.CVPR 2023
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 120 citations
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah et al.CVPR 2025
