Cross-Modal Retrieval with Heterogeneous Graph Embedding
Dapeng Chen, Min Wang, Haobin Chen, Lin Wu, Jing Qin, Wei Peng
Abstract
Conventional methods address the cross-modal retrieval problem by projecting the multi-modal data into a shared representation space. Such a strategy will inevitably lose the modality-specific information, leading to decreased retrieval accuracy. In this paper, we propose heterogeneous graph embeddings to preserve more abundant cross-modal information. The embedding from one modality will be compensated with the aggregated embeddings from the other modality. In particular, a self-denoising tree search is designed to reduce the "label noise" problem, making the heterogeneous neighborhood more semantically relevant. The dual-path aggregation tackles the "modality imbalance" problem, giving each sample comprehensive dual-modality information. The final heterogeneous graph embedding is obtained by feeding the aggregated dual-modality features to the cross-modal self-attention module. Experiments conducted on cross-modality person re-identification and image-text retrieval task validate the superiority and generality of the proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 00a84862-b9b0-4e61-a203-fed1db72496fCited by top-tier papers4
- Seq-HGNN: Learning Sequential Node Representation on Heterogeneous GraphChenguang Du, Kaichun Yao, Hengshu Zhu, Deqing Wang et al.SIGIR 2023 · 7 citations
- PRISM: Partial-label Relational Inference with Spatial and Spectral CuesYiyang Gu, Wenrui Wu, Yifang Qin, Taian Guo et al.ICLR 2026
- Hierarchical Encoding Tree with Modality Mixup for Cross-modal HashingZhiping Xiao, Junyu Luo, Hang Zhou, Yusheng Zhao et al.ICLR 2026
- Scalable Semi-supervised Community Search via Graph Transformer on Attributed Heterogeneous Information NetworksLinlin Ding, Zhaosong Zhao, Mo Li, Yishan Pan et al.AAAI 2026
Related papers
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang et al.ACM MM 2021 · 14 citations
- Dual Adversarial Graph Neural Networks for Multi-label Cross-modal RetrievalShengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang et al.AAAI 2021 · 72 citations
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 20 citations
- HL-CMR: Hypergraph Learning for Cross-Modal RetrievalGuohui Ding, Jing Li, Yimin Xu, Rui ZhouWWW 2026
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li et al.ICDE 2024 · 14 citations
