Hierarchical Encoding Tree with Modality Mixup for Cross-modal Hashing
Zhiping Xiao, Junyu Luo, Hang Zhou, Yusheng Zhao, Xiao Luo, Pengyun Wang, Wei Ju, Siyu Heng, Ming Zhang
摘要
Cross-modal retrieval is a fundamental task that aims to learn semantic correspondences across different data modalities, such as visual and textual modalities. Unsupervised hashing methods can efficiently manage large-scale data and can be effectively applied to cross-modal retrieval studies. However, existing methods typically fail to fully exploit the hierarchical semantic structure within text and image data, where instances naturally organize into multi-level communities of varying granularity. Moreover, the commonly-used direct modal alignment cannot effectively bridge the semantic gap between these two modalities. To address these issues, we introduce a novel Hierarchical Encoding Tree with Modality Mixup (HINT) method, which achieves effective cross-modal retrieval by extracting hierarchical cross-modal relations. HINT constructs a cross-modal encoding tree guided by hierarchical structural entropy and generates proxy samples of text and image modalities for each instance from the encoding tree. Through the curriculum-based mixup of proxy samples, HINT achieves progressive modal alignment and effective cross-modal retrieval. We also conduct cross-modal consistency learning to achieve global-view semantic alignment between text and image representations. Extensive experiments on a range of cross-modal retrieval datasets demonstrate the superiority of HINT over state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
- Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node EmbeddingsYu Chen, Lingfei Wu, Mohammed J. ZakiNeurIPS 2020 · 被引用 559 次
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding 等NeurIPS 2021 · 被引用 215 次
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingJun Yu, Hao Zhou, Yibing Zhan, Dacheng TaoAAAI 2021 · 被引用 184 次
相关 Paper
- Vision-guided Text Mining for Unsupervised Cross-modal Hashing with Community Similarity QuantizationHaozhi Fan, Yuan CaoAAAI 2025 · 被引用 9 次
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
- Unifying Two-Stream Encoders with Transformers for Cross-Modal RetrievalYi Bin, Haoxuan Li, Yahui Xu, Xing Xu 等ACM MM 2023 · 被引用 33 次
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing 等AAAI 2024 · 被引用 24 次
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 被引用 2 次
