Stationary and Clustering Transformer Hashing for Cross-modal Retrieval
Zhan Yang, Yiran Liu, Youyuan Huang, Yinan Li
Abstract
Unsupervised cross-modal hashing has gained significant attention for efficient retrieval between heterogeneous modalities through encoding data into the unified binary representations, offering low storage cost and fast response. However, the constraints of existing methods persist in bridging the cross-modal semantic gap and capturing fine-grained global semantic structures without explicit labels. In this paper, we propose an innovative unsupervised Stationary distribution and soft Clustering Transformer Hashing approach for cross-modal retrieval, denoted as SCTH. Initially, a Transformer-based modality fusion encoder is employed to extract abundant cross-modal semantic representations, further integrated with contrastive hashing to minimize the semantic gap. To enhance the inter-modal alignment, a pseudo-classifier clustering module with entropy-regularized contrastive loss is presented, ensuring balanced and diverse cluster assignments in unsupervised settings. Additionally, a Markovian stationary distribution strategy stabilizes the feature representations through mitigating the interference of noise and outliers. Comprehensive experiments on MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets validate that SCTH outperforms state-of-the-art hashing methods in cross-modal retrieval tasks, demonstrating superior generalization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e611d00-3cbb-415b-978f-24d06749f691Builds on14
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingJun Yu, Hao Zhou, Yibing Zhan, Dacheng TaoAAAI 2021 · 184 citations
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang et al.ACM MM 2023 · 44 citations
- Prototype-guided Knowledge Transfer for Federated Unsupervised Cross-modal HashingJingzhi Li, Fengling Li, Lei Zhu, Hui Cui et al.ACM MM 2023 · 33 citations
- OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask LearningSiddharth Srivastava, Gaurav SharmaCVPR 2024 · 32 citations
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren et al.AAAI 2025 · 25 citations
Related papers
- Unsupervised Similarity-Fusion Transformer Hashing for Multimodal RetrievalZhan Yang, Binghong Chen, Jiajun Tang, Yinan LiACM MM 2025
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan et al.ACM MM 2024 · 1 citation
- Bit-aware Semantic Transformer Hashing for Multi-modal RetrievalWentao Tan, Lei Zhu, Weili Guan, Jingjing Li et al.SIGIR 2022 · 33 citations
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 9 citations
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
