Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image Retrieval
Gangjian Zhang, Shikui Wei, Huaxin Pang, Yao Zhao
摘要
Composed image retrieval aims at performing image retrieval task by giving a reference image and a complementary text piece. Since composing both image and text information can accurately model the users' search intent, composed image retrieval can perform target-specific image retrieval task and be potentially applied to many scenarios such as interactive product search. However, two key challenging issues must be addressed in composed image retrieval occasion. One of them is how to fuse heterogeneous image and text piece in the query into a complementary feature space. The other is how to bridge the heterogeneous gap between text pieces in the query and images in the database. To address the issues, we propose an end-to-end framework for composed image retrieval, which consists of three key components including Multi-modal Complementary Fusion (MCF), Cross-modal Guided Pooling (CGP), and Relative Caption-aware Consistency (RCC). By incorporating MCF and CGP modules, we can fully integrate the complementary information of image and text piece in the query through multiple deep interactions and aggregate obtained local features into an embedding vector. To bridge the heterogeneous gap, we introduce the RCC constraint to align text pieces in the query and images in the database. Extensive experiments on four public benchmark datasets show that the proposed composed image retrieval framework achieves outstanding performance against the state-of-the-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei 等ACM MM 2023 · 被引用 53 次
- Aligning Composed Query with Image via Discriminative Perception from Negative CorrespondencesYifan Wang, Wuliang Huang, Chun YuanAAAI 2025 · 被引用 3 次
- SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image RetrievalYi Sun, Jinyu Xu, Qing Xie, Jiachen Li 等WWW 2026 · 被引用 1 次
相关 Paper
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 29 次
- Comprehensive Relationship Reasoning for Composed Query Based Image RetrievalFeifei Zhang, Ming Yan, Ji Zhang, Changsheng XuACM MM 2022 · 被引用 21 次
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou 等CVPR 2025
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 被引用 39 次
- CaLa: Complementary Association Learning for Augmenting Comoposed Image RetrievalXintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu 等SIGIR 2024 · 被引用 13 次
