Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image Retrieval
Feifei Zhang, Mingliang Xu, Qirong Mao, Changsheng Xu
Abstract
Cross-model retrieval has attracted much attention in recent years due to its wide applications. Conventional approaches usually take one modality as query to retrieve relevant data of another modality. In this paper, we devote to an emerging task in cross-modal retrieval, Composing Text and Image to Image Retrieval (CTI-IR), which aims at retrieving images relevant to a query image with text describing desired modifications to the query image. Compared with conventional cross-modal retrieval, the new task is particularly useful for the retrieval that the query image does not perfectly match the user's expectations. Generally, the CTI-IR involves two underlying problems: how to manipulate visual features of the query image specified by the text, and how to model the modality gap between the query and target. Most previous methods focus on solving the second problem. In this paper, we aim to deal with both problems simultaneously in a unified model. Specifically, the proposed method is based on the graph attention network and adversarial learning network, which enjoys several merits. First, the query image and the modification text are constructed in a relation graph for learning text-adaptive representations. Second, semantic contents from the text are injected into the visual features through graph attention. Third, an adversarial loss is incorporated into the conventional cross-modal retrieval loss to learn more discriminative modality invariant representations for CTI-IR. Extensive experiments on three benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7973b927-ff7b-4160-9728-904cb4eccbb0Cited by top-tier papers8
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- Progressive Learning for Image Retrieval with Hybrid-Modality QueriesYida Zhao, Yuqing Song, Qin JinSIGIR 2022 · 35 citations
- Dynamic Weighted Combiner for Mixed-Modal Image RetrievalFuxiang Huang, Lei Zhang, Xiaowei Fu, Suqi SongAAAI 2024 · 28 citations
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang et al.NeurIPS 2024 · 18 citations
Related papers
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 29 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalFeifei Zhang, Ming Yan, Ji Zhang, Changsheng XuACM MM 2022 · 21 citations
- CaLa: Complementary Association Learning for Augmenting Comoposed Image RetrievalXintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu et al.SIGIR 2024 · 13 citations
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
