Polysemic Semantic Instance Network for Cross-Modal Hashing
Shuo Han, Qibing Qin, Kezhen Xie, Wenfeng Zhang, Lei Huang
Abstract
Hashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performance. Most existing methods rely on one-to-one mappings that preserve only global average semantics, which fail to capture the intrinsic polysemous structures embedded within individual samples. To address this issue, we propose a novel Deep Polysemic Semantic Instance Hashing (DPSIH) method and design a Diverse Semantic Instance Embedding (DSIE) module. This module integrates local and global features through multi-head self-attention and residual learning, generating multiple diverse embeddings per sample to effectively capture fine-grained and polysemous semantic structures. Furthermore, we design a multi-embedding semantic correlation constraint that relaxes strict alignment restrictions to improve robustness under partial alignment, and introduce Maximum Mean Discrepancy (MMD) regularization to alleviate cross-modal distribution shifts. Additionally, an embedding diversity mechanism is proposed to prevent all embeddings from collapsing into a central or averaged representation, thereby enhancing semantic diversity. Extensive experiments on four benchmark datasets demonstrate that DPSIH significantly outperforms state-of-the-art methods and effectively improves the modeling of semantic ambiguity in cross-modal retrieval tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57314dde-7d0d-4b3f-adac-bccef26c9e5dBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 362 citations
- Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingJun Yu, Hao Zhou, Yibing Zhan, Dacheng TaoAAAI 2021 · 184 citations
- Differentiable Cross-modal Hashing via Multimodal TransformersJunfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong et al.ACM MM 2022 · 93 citations
- Dual Adversarial Graph Neural Networks for Multi-label Cross-modal RetrievalShengsheng Qian, Dizhan Xue, Huaiwen Zhang, Quan Fang et al.AAAI 2021 · 72 citations
Related papers
- Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal RetrievalKun Cheng, Qibing Qin, Wenfeng Zhang, Lei Huang et al.ACM MM 2025 · 4 citations
- Dual Self-Paced Cross-Modal HashingYuan Sun, Jian Dai, Zhenwen Ren, Yingke Chen et al.AAAI 2024 · 35 citations
- Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalYishu Liu, Qingpeng Wu, Zheng Zhang, Jingyi Zhang et al.ACM MM 2023 · 44 citations
- Distribution Consistency Guided Hashing for Cross-Modal RetrievalYuan Sun, Kaiming Liu, Yongxiang Li, Zhenwen Ren et al.ACM MM 2024 · 11 citations
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan et al.SIGIR 2020 · 214 citations
