Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image Retrieval
Feifei Zhang, Mingliang Xu, Qirong Mao, Changsheng Xu
摘要
Cross-model retrieval has attracted much attention in recent years due to its wide applications. Conventional approaches usually take one modality as query to retrieve relevant data of another modality. In this paper, we devote to an emerging task in cross-modal retrieval, Composing Text and Image to Image Retrieval (CTI-IR), which aims at retrieving images relevant to a query image with text describing desired modifications to the query image. Compared with conventional cross-modal retrieval, the new task is particularly useful for the retrieval that the query image does not perfectly match the user's expectations. Generally, the CTI-IR involves two underlying problems: how to manipulate visual features of the query image specified by the text, and how to model the modality gap between the query and target. Most previous methods focus on solving the second problem. In this paper, we aim to deal with both problems simultaneously in a unified model. Specifically, the proposed method is based on the graph attention network and adversarial learning network, which enjoys several merits. First, the query image and the modification text are constructed in a relation graph for learning text-adaptive representations. Second, semantic contents from the text are injected into the visual features through graph attention. Third, an adversarial loss is incorporated into the conventional cross-modal retrieval loss to learn more discriminative modality invariant representations for CTI-IR. Extensive experiments on three benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan 等SIGIR 2021 · 被引用 72 次
- Progressive Learning for Image Retrieval with Hybrid-Modality QueriesYida Zhao, Yuqing Song, Qin JinSIGIR 2022 · 被引用 35 次
- Dynamic Weighted Combiner for Mixed-Modal Image RetrievalFuxiang Huang, Lei Zhang, Xiaowei Fu, Suqi SongAAAI 2024 · 被引用 28 次
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang 等ACM MM 2022 · 被引用 25 次
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang 等NeurIPS 2024 · 被引用 18 次
相关 Paper
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 29 次
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei 等ACM MM 2023 · 被引用 53 次
- Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalFeifei Zhang, Ming Yan, Ji Zhang, Changsheng XuACM MM 2022 · 被引用 21 次
- CaLa: Complementary Association Learning for Augmenting Comoposed Image RetrievalXintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu 等SIGIR 2024 · 被引用 13 次
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou 等CVPR 2025
