Lune

NeurIPS2025Top-tier venue

Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval

Shuang Zhang, Yue Wu, Lei Shi, Yingxue Zhang, Feifei Kou, Huilong Jin, Pengfei Zhang, Meiyu Liang, Mingying Xu

2025Year
2Citations
2Top-tier citations

Abstract

The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.

between different data types. In recent years, cross-modal hashing retrieval [4][5][6][7] has attracted widespread attention due to its advantages of fast retrieval and efficient storage. It uses hashing technology to convert high-dimensional data into low-dimensional binary hash codes, thereby reducing computational complexity and storage requirements while retaining semantic information.

Currently, some scholars have proposed a variety of new cross-modal hashing retrieval methods. Neural network technologies, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have been widely used in cross-modal hashing retrieval. CNN can effectively extract semantic information from images through its powerful feature extraction capabilities, while GAN generates robust hash codes through adversarial training of generators and discriminators, thereby improving the performance of cross-modal retrieval. In addition, large language models (LLMs) have also been introduced into cross-modal hashing retrieval, which enhance the semantic representation of text modalities through their powerful natural language processing capabilities, thereby improving the accuracy of cross-modal matching.

Although many methods have achieved good results in the field of cross-modal hashing retrieval, they still face some challenges. Due to the huge semantic gap between different modalities [8][9] often leads to inconsistent cross-modal representations, many noncritical information or noise[10] may affect the matching accuracy, resulting in similar images and texts being mismatched [11]. Secondly, when processing features, traditional hash layers often ignore the importance differences between different channels [12], which can lead to insufficient capture of key information and difficulty in effectively suppressing noise and redundant information. In addition, when hash codes are generated, they usually rely on a single optimization goal, which may lead to insufficient performance of hash codes in cross-modal matching.

To effectively address these challenges, we proposed a method called auxiliary hashing learning (AHLR). It significantly improves feature extraction and alignment capability by introducing a dynamic mask mechanism. Specifically, the dynamic mask can automatically identify and weight key information in the image and text during the training process, effectively improving the accuracy of matching of cross-modal features. In addition, we also constructed an auxiliary hashing layer that can adaptively weight the features of each channel, thereby solving the problem of channel information imbalance, while enhancing the ability to capture key information and effectively suppressing noise interference. Finally, by introducing a contrastive loss function, minimizing the distance between similar samples, and maximizing the distance between heterogeneous samples, the distinguishing ability of hash codes in cross-modal retrieval is effectively enhanced, thereby improving the retrieval accuracy. The main contributions of this paper are as follows:

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 322727f7-4a32-4689-b6eb-1c469f4e278c

Cited by top-tier papers2

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines