Stationary and Clustering Transformer Hashing for Cross-modal Retrieval
Zhan Yang, Yiran Liu, Youyuan Huang, Yinan Li
摘要
Unsupervised cross-modal hashing has gained significant attention for efficient retrieval between heterogeneous modalities through encoding data into the unified binary representations, offering low storage cost and fast response. However, the constraints of existing methods persist in bridging the cross-modal semantic gap and capturing fine-grained global semantic structures without explicit labels. In this paper, we propose an innovative unsupervised Stationary distribution and soft Clustering Transformer Hashing approach for cross-modal retrieval, denoted as SCTH. Initially, a Transformer-based modality fusion encoder is employed to extract abundant cross-modal semantic representations, further integrated with contrastive hashing to minimize the semantic gap. To enhance the inter-modal alignment, a pseudo-classifier clustering module with entropy-regularized contrastive loss is presented, ensuring balanced and diverse cluster assignments in unsupervised settings. Additionally, a Markovian stationary distribution strategy stabilizes the feature representations through mitigating the interference of noise and outliers. Comprehensive experiments on MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets validate that SCTH outperforms state-of-the-art hashing methods in cross-modal retrieval tasks, demonstrating superior generalization performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingJun Yu, Hao Zhou, Yibing Zhan, Dacheng TaoAAAI 2021 · 被引用 184 次
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang 等ACM MM 2023 · 被引用 44 次
- Prototype-guided Knowledge Transfer for Federated Unsupervised Cross-modal HashingJingzhi Li, Fengling Li, Lei Zhu, Hui Cui 等ACM MM 2023 · 被引用 33 次
- OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask LearningSiddharth Srivastava, Gaurav SharmaCVPR 2024 · 被引用 32 次
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren 等AAAI 2025 · 被引用 25 次
相关 Paper
- Unsupervised Similarity-Fusion Transformer Hashing for Multimodal RetrievalZhan Yang, Binghong Chen, Jiajun Tang, Yinan LiACM MM 2025
- Similarity Preserving Transformer Cross-Modal Hashing for Video-Text RetrievalQianxin Huang, Siyao Peng, Xiaobo Shen, Yunhao Yuan 等ACM MM 2024 · 被引用 1 次
- Bit-aware Semantic Transformer Hashing for Multi-modal RetrievalWentao Tan, Lei Zhu, Weili Guan, Jingjing Li 等SIGIR 2022 · 被引用 33 次
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
