Robust Image Hashing Based on Contrastive Masked Autoencoder with Weak-Strong Augmentation Alignment
Cundian Yang, Guibo Luo, Yuesheng Zhu, Jiaqi Li, Xiyao Liu
Abstract
Recently, numerous robust image hashing schemes have been developed for content identification. However, many of these schemes face the challenges of maintaining discrimination while simultaneously resisting large-scale attacks. In this paper, we propose a robust image hashing scheme based on Contrastive Masked Autoencoder with weak-strong augmentation Alignment (CMAA). Leveraging contrastive learning, CMAA is designed to learn features that are robust to large-scale and hybrid attacks while maintaining the discrimination of those features. Specifically, it utilizes distribution divergence to align weak attack augmented features with strong attack augmented features, namely weak-strong augmentation alignment, to enhance the robustness to strong attacks. In addition, a masked vision transformer is incorporated to further enhance content identification performance. CMAA also includes a parameter-free quantization layer to mitigate the loss induced by binarization. Experimental results demonstrate that our method exhibits remarkable robustness against various attacks, including challenging ones such as rotation and hybrid attacks, and delivers excellent identification performance with a F1 score close to 1.0. Our code and supplementary materials are available on Github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Deep Unsupervised Image Hashing by Maximizing Bit EntropyYunqiang Li, Jan van GemertAAAI 2021 · 109 citations
- Rethinking the Augmentation Module in Contrastive Learning: Learning Hierarchical Augmentation Invariance with Expanded ViewsJunbo Zhang, Kaisheng MaCVPR 2022 · 39 citations
Related papers
- Spectral-Adaptive Adversarial Hashing for Robust Image RetrievalGang Zhou, Shibiao Xu, Xiaolong Zheng, Daniel Dajun ZengSIGIR 2026
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao et al.CVPR 2026
- When Perceptual Authentication Hashing Meets Neural Architecture SearchYuanding Zhou, Xinran Li, Yaodong Fang, Chuan QinACM MM 2023 · 3 citations
- Medical Vision-Language Pre-training with Multimodal Variational Masked Autoencoder for Robust Medical VQADexuan Xu, Yanyuan Chen, Yu Huang, Shihao E et al.ACM MM 2025
- Two-Stage Adversarial Training for Deep Hashing via Representation DistillationFei Zhu, Huashan Chen, Wanqian Zhang, Lin Wang et al.SIGIR 2025 · 2 citations
