Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval
Shuang Zhang, Yue Wu, Lei Shi, Yingxue Zhang, Feifei Kou, Huilong Jin, Pengfei Zhang, Meiyu Liang, Mingying Xu
摘要
The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.
between different data types. In recent years, cross-modal hashing retrieval [4][5][6][7] has attracted widespread attention due to its advantages of fast retrieval and efficient storage. It uses hashing technology to convert high-dimensional data into low-dimensional binary hash codes, thereby reducing computational complexity and storage requirements while retaining semantic information.
Currently, some scholars have proposed a variety of new cross-modal hashing retrieval methods. Neural network technologies, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have been widely used in cross-modal hashing retrieval. CNN can effectively extract semantic information from images through its powerful feature extraction capabilities, while GAN generates robust hash codes through adversarial training of generators and discriminators, thereby improving the performance of cross-modal retrieval. In addition, large language models (LLMs) have also been introduced into cross-modal hashing retrieval, which enhance the semantic representation of text modalities through their powerful natural language processing capabilities, thereby improving the accuracy of cross-modal matching.
Although many methods have achieved good results in the field of cross-modal hashing retrieval, they still face some challenges. Due to the huge semantic gap between different modalities [8][9] often leads to inconsistent cross-modal representations, many noncritical information or noise[10] may affect the matching accuracy, resulting in similar images and texts being mismatched [11]. Secondly, when processing features, traditional hash layers often ignore the importance differences between different channels [12], which can lead to insufficient capture of key information and difficulty in effectively suppressing noise and redundant information. In addition, when hash codes are generated, they usually rely on a single optimization goal, which may lead to insufficient performance of hash codes in cross-modal matching.
To effectively address these challenges, we proposed a method called auxiliary hashing learning (AHLR). It significantly improves feature extraction and alignment capability by introducing a dynamic mask mechanism. Specifically, the dynamic mask can automatically identify and weight key information in the image and text during the training process, effectively improving the accuracy of matching of cross-modal features. In addition, we also constructed an auxiliary hashing layer that can adaptively weight the features of each channel, thereby solving the problem of channel information imbalance, while enhancing the ability to capture key information and effectively suppressing noise interference. Finally, by introducing a contrastive loss function, minimizing the distance between similar samples, and maximizing the distance between heterogeneous samples, the distinguishing ability of hash codes in cross-modal retrieval is effectively enhanced, thereby improving the retrieval accuracy. The main contributions of this paper are as follows:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao 等CVPR 2026
- Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy LabelsXincheng Sun, Ruitao Pu, Guangsi Shi, Zhenwen Ren 等ICML 2026
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 被引用 261 次
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan 等SIGIR 2020 · 被引用 214 次
- Differentiable Cross-modal Hashing via Multimodal TransformersJunfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong 等ACM MM 2022 · 被引用 93 次
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren 等AAAI 2025 · 被引用 25 次
相关 Paper
- Adaptive Graph Attention Based Discrete Hashing for Incomplete Cross-modal RetrievalShuang Zhang, Yue Wu, Lei Shi, Huilong Jin 等AAAI 2026
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
- Alleviating the Inconsistency of Multimodal Data in Cross-Modal RetrievalTieying Li, Xiaochun Yang, Yiping Ke, Bin Wang 等ICDE 2024 · 被引用 8 次
- Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model for Incomplete Multimodal RetrievalYang Yu, Meiyu Liang, Wei Huang, Juncheng Zheng 等ACM MM 2025
- PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalQiang Zou, Shuli Cheng, Jiayi ChenCVPR 2025
