Lune

NeurIPS2025顶会

Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval

Shuang Zhang, Yue Wu, Lei Shi, Yingxue Zhang, Feifei Kou, Huilong Jin, Pengfei Zhang, Meiyu Liang, Mingying Xu

2025年份
2被引次数
2顶会引用

摘要

The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.

between different data types. In recent years, cross-modal hashing retrieval [4][5][6][7] has attracted widespread attention due to its advantages of fast retrieval and efficient storage. It uses hashing technology to convert high-dimensional data into low-dimensional binary hash codes, thereby reducing computational complexity and storage requirements while retaining semantic information.

Currently, some scholars have proposed a variety of new cross-modal hashing retrieval methods. Neural network technologies, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have been widely used in cross-modal hashing retrieval. CNN can effectively extract semantic information from images through its powerful feature extraction capabilities, while GAN generates robust hash codes through adversarial training of generators and discriminators, thereby improving the performance of cross-modal retrieval. In addition, large language models (LLMs) have also been introduced into cross-modal hashing retrieval, which enhance the semantic representation of text modalities through their powerful natural language processing capabilities, thereby improving the accuracy of cross-modal matching.

Although many methods have achieved good results in the field of cross-modal hashing retrieval, they still face some challenges. Due to the huge semantic gap between different modalities [8][9] often leads to inconsistent cross-modal representations, many noncritical information or noise[10] may affect the matching accuracy, resulting in similar images and texts being mismatched [11]. Secondly, when processing features, traditional hash layers often ignore the importance differences between different channels [12], which can lead to insufficient capture of key information and difficulty in effectively suppressing noise and redundant information. In addition, when hash codes are generated, they usually rely on a single optimization goal, which may lead to insufficient performance of hash codes in cross-modal matching.

To effectively address these challenges, we proposed a method called auxiliary hashing learning (AHLR). It significantly improves feature extraction and alignment capability by introducing a dynamic mask mechanism. Specifically, the dynamic mask can automatically identify and weight key information in the image and text during the training process, effectively improving the accuracy of matching of cross-modal features. In addition, we also constructed an auxiliary hashing layer that can adaptively weight the features of each channel, thereby solving the problem of channel information imbalance, while enhancing the ability to capture key information and effectively suppressing noise interference. Finally, by introducing a contrastive loss function, minimizing the distance between similar samples, and maximizing the distance between heterogeneous samples, the distinguishing ability of hash codes in cross-modal retrieval is effectively enhanced, thereby improving the retrieval accuracy. The main contributions of this paper are as follows:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖