Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search
Meiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing, Wei Huang, Zhe Xue
Abstract
Deep cross-modal hashing technology provides an effective and efficient cross-modal unified representation learning solution for cross-modal search. However, the existing methods neglect the implicit fine-grained multimodal knowledge relations between different modalities such as when the image contains information that is not directly described in the text. To tackle this problem, we propose a novel selfsupervised multi-grained multi-modal knowledge graph contrastive hashing method for cross-modal search (CMGCH). Firstly, in order to capture implicit fine-grained cross-modal semantic associations, a multi-modal knowledge graph is constructed, which represents the implicit multimodal knowledge relations between the image and text as inter-modal and intra-modal semantic associations. Secondly, a cross-modal graph contrastive attention network is proposed to reason on the multi-modal knowledge graph to sufficiently learn the implicit fine-grained inter-modal and intra-modal knowledge relations. Thirdly, a cross-modal multi-granularity contrastive embedding learning mechanism is proposed, which fuses the global coarse-grained and local fine-grained embeddings by multihead attention mechanism for inter-modal and intra-modal contrastive learning, so as to enhance the crossmodal unified representations with stronger discriminativeness and semantic consistency preserving power. With the joint training of intra-modal and inter-modal contrast, the invariant and modal-specific information of different modalities can be maintained in the final cross-modal unified hash space. Extensive experiments on several cross-modal benchmark datasets demonstrate that the proposed CMGCH outperforms the state-of the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- HLMEA: Unsupervised Entity Alignment Based on Hybrid Language ModelsXiongnan Jin, Zhilin Wang, Jinpeng Chen, Liu Yang et al.AAAI 2025 · 5 citations
- Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal RetrievalShuang Zhang, Yue Wu, Lei Shi, Yingxue Zhang et al.NeurIPS 2025 · 2 citations
- Hierarchical Encoding Tree with Modality Mixup for Cross-modal HashingZhiping Xiao, Junyu Luo, Hang Zhou, Yusheng Zhao et al.ICLR 2026
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao et al.CVPR 2026
- HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph CompletionDi Wang, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Semantic Structure Enhanced Contrastive Adversarial Hash Network for Cross-media Representation LearningMeiYu Liang, Junping Du, Xiaowen Cao, Yang Yu et al.ACM MM 2022 · 15 citations
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang et al.ACM MM 2023 · 44 citations
- An End-To-End Graph Attention Network Hashing for Cross-Modal RetrievalHuilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang et al.NeurIPS 2024 · 18 citations
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- Unsupervised Multimodal Graph Contrastive Semantic Anchor Space Dynamic Knowledge Distillation Network for Cross-Media Hash RetrievalYang Yu, Meiyu Liang, Mengran Yin, Kangkang Lu et al.ICDE 2024 · 5 citations
