Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval
Shuang Zhang, Yue Wu, Lei Shi, Yingxue Zhang, Feifei Kou, Huilong Jin, Pengfei Zhang, Meiyu Liang, Mingying Xu
Abstract
The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.
between different data types. In recent years, cross-modal hashing retrieval [4][5][6][7] has attracted widespread attention due to its advantages of fast retrieval and efficient storage. It uses hashing technology to convert high-dimensional data into low-dimensional binary hash codes, thereby reducing computational complexity and storage requirements while retaining semantic information.
Currently, some scholars have proposed a variety of new cross-modal hashing retrieval methods. Neural network technologies, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have been widely used in cross-modal hashing retrieval. CNN can effectively extract semantic information from images through its powerful feature extraction capabilities, while GAN generates robust hash codes through adversarial training of generators and discriminators, thereby improving the performance of cross-modal retrieval. In addition, large language models (LLMs) have also been introduced into cross-modal hashing retrieval, which enhance the semantic representation of text modalities through their powerful natural language processing capabilities, thereby improving the accuracy of cross-modal matching.
Although many methods have achieved good results in the field of cross-modal hashing retrieval, they still face some challenges. Due to the huge semantic gap between different modalities [8][9] often leads to inconsistent cross-modal representations, many noncritical information or noise[10] may affect the matching accuracy, resulting in similar images and texts being mismatched [11]. Secondly, when processing features, traditional hash layers often ignore the importance differences between different channels [12], which can lead to insufficient capture of key information and difficulty in effectively suppressing noise and redundant information. In addition, when hash codes are generated, they usually rely on a single optimization goal, which may lead to insufficient performance of hash codes in cross-modal matching.
To effectively address these challenges, we proposed a method called auxiliary hashing learning (AHLR). It significantly improves feature extraction and alignment capability by introducing a dynamic mask mechanism. Specifically, the dynamic mask can automatically identify and weight key information in the image and text during the training process, effectively improving the accuracy of matching of cross-modal features. In addition, we also constructed an auxiliary hashing layer that can adaptively weight the features of each channel, thereby solving the problem of channel information imbalance, while enhancing the ability to capture key information and effectively suppressing noise interference. Finally, by introducing a contrastive loss function, minimizing the distance between similar samples, and maximizing the distance between heterogeneous samples, the distinguishing ability of hash codes in cross-modal retrieval is effectively enhanced, thereby improving the retrieval accuracy. The main contributions of this paper are as follows:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 322727f7-4a32-4689-b6eb-1c469f4e278cCited by top-tier papers2
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao et al.CVPR 2026
- Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy LabelsXincheng Sun, Ruitao Pu, Guangsi Shi, Zhenwen Ren et al.ICML 2026
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan et al.SIGIR 2020 · 214 citations
- Differentiable Cross-modal Hashing via Multimodal TransformersJunfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong et al.ACM MM 2022 · 93 citations
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren et al.AAAI 2025 · 25 citations
Related papers
- Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal RetrievalShuang Zhang, Yue Wu, Lei Shi, Huilong Jin et al.AAAI 2026
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang et al.ICDE 2024 · 8 citations
- Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model for Incomplete Multimodal RetrievalYang Yu, Meiyu Liang, Wei Huang, Juncheng Zheng et al.ACM MM 2025
- PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalQiang Zou, Shuli Cheng, Jiayi ChenCVPR 2025
