Image Search with Text Feedback by Deep Hierarchical Attention Mutual Information Maximization
Chunbin Gu, Jiajun Bu, Zhen Zhang, Zhi Yu, Dongfang Ma, Wei Wang
Abstract
Image retrieval with text feedback is an emerging research topic with the objective of integrating inputs from multiple modalities as queries. In this paper, queries contain a reference image plus text feedback that describes modifications between this image and the desired image. The existing work for this task mainly focuses on designing a new fusion network to compose the image and text. Still, little research pays attention to the modality gap caused by the inconsistent distribution of features from different modalities, which dramatically influences the feature fusion and similarity learning between queries and the desired image. We propose a Distribution-Aligned Text-based Image Retrieval (DATIR) model, which consists of attention mutual information maximization and hierarchical mutual information maximization, to bridge this gap by increasing non-linear statistic dependencies between representations of different modalities. More specifically, attention mutual information maximization narrows the modality gap between different input modalities by maximizing mutual information between the text representation and its semantically consistent representation captured from the reference image and the desired image by the difference transformer. For hierarchical mutual information maximization, it aligns distributions of features from the image modality and the fusion modality by estimating mutual information between a single-layer representation in the fusion network and the multi-level representations in the desired image encoder. Extensive experiments on three large-scale benchmark datasets demonstrate that we can bridge the modality gap between different modalities and achieve state-of-the-art retrieval performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu et al.AAAI 2025 · 59 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Progressive Learning for Image Retrieval with Hybrid-Modality QueriesYida Zhao, Yuqing Song, Qin JinSIGIR 2022 · 35 citations
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu et al.SIGIR 2024 · 29 citations
- Dynamic Weighted Combiner for Mixed-Modal Image RetrievalFuxiang Huang, Lei Zhang, Xiaowei Fu, Suqi SongAAAI 2024 · 28 citations
Related papers
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu et al.ACM MM 2023 · 33 citations
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 29 citations
- Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image RetrievalGangjian Zhang, Shikui Wei, Huaxin Pang, Yao ZhaoACM MM 2021 · 34 citations
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li et al.ICDE 2024 · 14 citations
