Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval
Shengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang, Changsheng Xu
摘要
Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, these methods typically suffer from following limitations: 1) They usually attempt to bridge the modality gap by designing losses in the common representation space which may not be sufficient to eliminate potential heterogeneity of different modalities in the common space. 2) They typically treat labels as independent individuals and ignore label relationships which are important for constructing semantic links between multimodal data. In this work, we propose a novel Dual Adversarial Graph Neural Networks (DAGNN) composed of the dual generative adversarial networks and the multi-hop graph neural networks, which learn modality-invariant and discriminative common representations for cross-modal retrieval. Firstly, we construct the dual generative adversarial networks to project multimodal data into a common representation space. Secondly, we leverage the multi-hop graph neural networks, in which a layer aggregation mechanism is proposed to exploit multi-hop propagation information, to capture the label correlation dependency and learn inter-dependent classifiers. Comprehensive experiments conducted on two cross-modal retrieval benchmark datasets, NUS-WIDE and MIRFlickr, indicate the superiority of DAGNN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen 等ICCV 2021 · 被引用 172 次
- Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalYizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi 等AAAI 2023 · 被引用 39 次
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing 等AAAI 2024 · 被引用 24 次
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 被引用 19 次
- Negative Pre-aware for Noisy Cross-Modal MatchingXu Zhang, Hao Li, Mang YeAAAI 2024 · 被引用 18 次
它引用的顶会 Paper3
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 被引用 261 次
- Unsupervised Domain Adaptive Graph Convolutional NetworksMan Wu, Shirui Pan, Chuan Zhou, Xiaojun Chang 等WWW 2020 · 被引用 221 次
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan 等SIGIR 2020 · 被引用 214 次
相关 Paper
- Exploring Graph-Structured Semantics for Cross-Modal RetrievalLei Zhang, Leiting Chen, Chuan Zhou, Fan Yang 等ACM MM 2021 · 被引用 14 次
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng 等AAAI 2025 · 被引用 1 次
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 被引用 233 次
- Cross-Modal Retrieval with Heterogeneous Graph EmbeddingDapeng Chen, Min Wang, Haobin Chen, Lin Wu 等ACM MM 2022 · 被引用 38 次
- Invariant Meets Specific: A Scalable Harmful Memes Detection FrameworkChuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin HuACM MM 2023 · 被引用 9 次
