CaLa: Complementary Association Learning for Augmenting Comoposed Image Retrieval
Xintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu, Bingwen Hu, Xueming Qian
Abstract
Composed image retrieval (CIR) is the task of searching target images using an image-text pair as a query. Given the straightforward relation of query pair-target image, the dominant methods follow the learning paradigm of common image-text retrieval and simply model this problem as the query-target matching problem. Particularly, the common practice first encodes the multi-modal query into one feature and then aligns it with the target image. However, such a learning paradigm only explores the naive relation in the triplets. We argue that CIR triplets encompass additional associations besides the primary query-target relation, which is overlooked in existing works. In this paper, we disclose two new relations residing in the triplets by viewing the triplet as a graph node. In analogy with the graph node, we mine two associations of text-bridged image alignment and complementary text reasoning. The text-bridged image alignment considers composed image retrieval as a specialized form of image retrieval, where the query text acts as a bridge between the query image and the target one, and a hinge-based cross attention is proposed to incorporate this relation into the network learning. On the other hand, the association of complementary text reasoning regards composed image retrieval as a specific type of cross-modal retrieval, where the composite two images are used to reason the complementary text. To integrate these views effectively, a twin attention-based compositor is designed. By combining these two types of complementary associations with the explicit query pair-target image relation, we establish a comprehensive set of constraints for composed image retrieval. With the above designs, we finally developed our CaLa, a Complementary Association Learning framework for Augmenting Composed Image Retrieval. Experimental evaluations are conducted on the widely-used CIRR and FashIionIQ benchmarks with multiple backbones to validate the effectiveness of our CaLa. The results demonstrate the superiority of our method in the composed image retrieval task. Our code and models are available at https://github.com/Chiangsonw/CaLa
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 033d05a0-cf01-4cd6-b523-8dd673bf6b75Cited by top-tier papers19
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang et al.CVPR 2026 · 16 citations
- Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image RetrievalZhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang et al.CVPR 2026 · 16 citations
- INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li et al.AAAI 2026 · 12 citations
- OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu et al.ACM MM 2025 · 10 citations
- HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang et al.AAAI 2026 · 8 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
Related papers
- Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalFeifei Zhang, Ming Yan, Ji Zhang, Changsheng XuACM MM 2022 · 21 citations
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang et al.CVPR 2026 · 3 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
