An End-To-End Graph Attention Network Hashing for Cross-Modal Retrieval
Huilong Jin, Yingxue Zhang, Lei Shi, Shuang Zhang, Feifei Kou, Jiapeng Yang, Chuangying Zhu, Jia Luo
Abstract
Due to its low storage cost and fast search speed, cross-modal retrieval based on hashing has attracted widespread attention and is widely used in real-world applications of social media search. However, most existing hashing methods are often limited by uncomprehensive feature representations and semantic associations, which greatly restricts their performance and applicability in practical applications. To deal with this challenge, in this paper, we propose an end-to-end graph attention network hashing (EGATH) for cross-modal retrieval, which can not only capture direct semantic associations between images and texts but also match semantic content between different modalities. We adopt the contrastive language image pretraining (CLIP) combined with the Transformer to improve understanding and generalization ability in semantic consistency across different data modalities. The classifier based on graph attention network is applied to obtain predicted labels to enhance cross-modal feature representation. We construct hash codes using an optimization strategy and loss function to preserve the semantic information and compactness of the hash code. Comprehensive experiments on the NUS-WIDE, MIRFlickr25K, and MS-COCO benchmark datasets show that our EGATH significantly outperforms against several state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 417c83af-30bb-4d8c-a98e-0aa6586bdb72Cited by top-tier papers4
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal RetrievalLanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu et al.NeurIPS 2025 · 12 citations
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
- Online Cross-Modal Hashing with Expanding Label SpaceWentao Fan, Chao Zhang, Chunlin Chen, Huaxiong LiAAAI 2026
- Meta-Guided Sample Reweighting for Robust Cross-Modal Hashing Retrieval with Noisy LabelsZiang Tan, Weitao An, Erkun YangAAAI 2026
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan et al.SIGIR 2020 · 214 citations
- Differentiable Cross-modal Hashing via Multimodal TransformersJunfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong et al.ACM MM 2022 · 93 citations
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
Related papers
- Lightweight Contrastive Distilled Hashing for Online Cross-modal RetrievalJiaxing Li, Lin Jiang, Zeqi Ma, Kaihang Jiang et al.AAAI 2025 · 4 citations
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal RetrievalShuang Zhang, Yue Wu, Lei Shi, Huilong Jin et al.AAAI 2026
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalQiang Zou, Shuli Cheng, Jiayi ChenCVPR 2025
