Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval
Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang, Changsheng Xu
Abstract
Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, these methods typically suffer from following limitations: 1) They usually attempt to bridge the modality gap by designing losses in the common representation space which may not be sufficient to eliminate potential heterogeneity of different modalities in the common space. 2) They typically treat labels as independent individuals and ignore label relationships which are important for constructing semantic links between multimodal data. In this work, we propose a novel Dual Adversarial Graph Neural Networks (DAGNN) composed of the dual generative adversarial networks and the multi-hop graph neural networks, which learn modality-invariant and discriminative common representations for cross-modal retrieval. Firstly, we construct the dual generative adversarial networks to project multimodal data into a common representation space. Secondly, we leverage the multi-hop graph neural networks, in which a layer aggregation mechanism is proposed to exploit multi-hop propagation information, to capture the label correlation dependency and learn inter-dependent classifiers. Comprehensive experiments conducted on two cross-modal retrieval benchmark datasets, NUS-WIDE and MIRFlickr, indicate the superiority of DAGNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2c4fa65-3ca0-4007-99a8-df86baba0e7aCited by top-tier papers9
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen et al.ICCV 2021 · 172 citations
- Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalYizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi et al.AAAI 2023 · 39 citations
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing et al.AAAI 2024 · 24 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- Negative Pre-aware for Noisy Cross-Modal MatchingXu Zhang, Hao Li, Mang YeAAAI 2024 · 18 citations
Builds on3
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Unsupervised Domain Adaptive Graph Convolutional NetworksMan Wu, Shirui Pan, Chuan Zhou, Xiaojun Chang et al.WWW 2020 · 221 citations
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan et al.SIGIR 2020 · 214 citations
Related papers
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang et al.ACM MM 2021 · 14 citations
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng et al.AAAI 2025 · 1 citation
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 233 citations
- Cross-Modal Retrieval with Heterogeneous Graph EmbeddingDapeng Chen, Min Wang, Haobin Chen, Lin Wu et al.ACM MM 2022 · 38 citations
- Invariant Meets Specific: A Scalable Harmful Memes Detection FrameworkChuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin HuACM MM 2023 · 9 citations
