Dual Compositional Learning in Interactive Image Retrieval
Jongseok Kim, Youngjae Yu, Hoeseong Kim, Gunhee Kim
Abstract
We present an approach named Dual Composition Network (DCNet) for interactive image retrieval that searches for the best target image for a natural language query and a reference image. To accomplish this task, existing methods have focused on learning a composite representation of the reference image and the text query to be as close to the embedding of the target image as possible. We refer this approach as Composition Network. In this work, we propose to close the loop with Correction Network that models the difference between the reference and target image in the embedding space and matches it with the embedding of the text query. That is, we consider two cyclic directional mappings for triplets of (reference image, text query, target image) by using both Composition Network and Correction Network. We also propose a joint training loss that can further improve the robustness of multimodal representation learning. We evaluate the proposed model on three benchmark datasets for multimodal retrieval: Fashion-IQ, Shoes, and Fashion200K. Our experiments show that our DCNet achieves new state-of-the-art performance on all three datasets, and the addition of Correction Network consistently improves multiple existing methods that are solely based on Composition Network. Moreover, an ensemble of our model won the first place in Fashion-IQ 2020 challenge held in a CVPR 2020 workshop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- Effective conditioned and composed image retrieval combining CLIP-based featuresAlberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del BimboCVPR 2022 · 139 citations
- CoVR: Learning Composed Video Retrieval from Web Video CaptionsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolAAAI 2024 · 81 citations
- Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty RegularizationYiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu et al.ICLR 2024 · 80 citations
- Sentence-level Prompts Benefit Composed Image RetrievalYang Bai, Xinxing Xu, Yong Liu, Salman Khan et al.ICLR 2024 · 75 citations
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
Builds on6
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Interactive Dual Generative Adversarial Networks for Image CaptioningJunhao Liu, Kai Wang, Chunpu Xu, Zhou Zhao et al.AAAI 2020 · 35 citations
- Towards Hands-Free Visual Dialog Interactive RecommendationTong Yu, Yilin Shen, Hongxia JinAAAI 2020 · 18 citations
- Image Search With Text Feedback by Visiolinguistic Attention LearningYanbei Chen, Shaogang Gong, Loris BazzaniCVPR 2020
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
Related papers
- Multi-Schema Proximity Network for Composed Image RetrievalJiangming Shi, Xiangbo Yin, Yeyun Chen, Yachao Zhang et al.ICCV 2025 · 1 citation
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah et al.CVPR 2025
- Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and NegativesZhangchi Feng, Richong Zhang, Zhijie NieACM MM 2024 · 14 citations
- INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li et al.AAAI 2026 · 12 citations
