Cross-Modal Retrieval with Heterogeneous Graph Embedding
Dapeng Chen, Min Wang, Haobin Chen, Lin Wu, Jing Qin, Wei Peng
摘要
Conventional methods address the cross-modal retrieval problem by projecting the multi-modal data into a shared representation space. Such a strategy will inevitably lose the modality-specific information, leading to decreased retrieval accuracy. In this paper, we propose heterogeneous graph embeddings to preserve more abundant cross-modal information. The embedding from one modality will be compensated with the aggregated embeddings from the other modality. In particular, a self-denoising tree search is designed to reduce the "label noise" problem, making the heterogeneous neighborhood more semantically relevant. The dual-path aggregation tackles the "modality imbalance" problem, giving each sample comprehensive dual-modality information. The final heterogeneous graph embedding is obtained by feeding the aggregated dual-modality features to the cross-modal self-attention module. Experiments conducted on cross-modality person re-identification and image-text retrieval task validate the superiority and generality of the proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Seq-HGNN: Learning Sequential Node Representation on Heterogeneous GraphChenguang Du, Kaichun Yao, Hengshu Zhu, Deqing Wang 等SIGIR 2023 · 被引用 7 次
- PRISM: Partial-label Relational Inference with Spatial and Spectral CuesYiyang Gu, Wenrui Wu, Yifang Qin, Taian Guo 等ICLR 2026
- Hierarchical Encoding Tree with Modality Mixup for Cross-modal HashingZhiping Xiao, Junyu Luo, Hang Zhou, Yusheng Zhao 等ICLR 2026
- Scalable Semi-supervised Community Search via Graph Transformer on Attributed Heterogeneous Information NetworksLinlin Ding, Zhaosong Zhao, Mo Li, Yishan Pan 等AAAI 2026
相关 Paper
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang 等ACM MM 2021 · 被引用 14 次
- Dual Adversarial Graph Neural Networks for Multi-label Cross-modal RetrievalShengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang 等AAAI 2021 · 被引用 72 次
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 被引用 20 次
- HL-CMR: Hypergraph Learning for Cross-Modal RetrievalGuohui Ding, Jing Li, Yimin Xu, Rui ZhouWWW 2026
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li 等ICDE 2024 · 被引用 14 次
