Unsupervised Multi-Index Semantic Hashing
Christian Hansen, Casper Hansen, Jakob Grue Simonsen, Stephen Alstrup, Christina Lioma
Abstract
Semantic hashing represents documents as compact binary vectors (hash codes) and allows both efficient and effective similarity search in large-scale information retrieval. The state of the art has primarily focused on learning hash codes that improve similarity search effectiveness, while assuming a brute-force linear scan strategy for searching over all the hash codes, even though much faster alternatives exist. One such alternative is multi-index hashing, an approach that constructs a smaller candidate set to search over, which depending on the distribution of the hash codes can lead to sub-linear search time. In this work, we propose Multi-Index Semantic Hashing (MISH), an unsupervised hashing model that learns hash codes that are both effective and highly efficient by being optimized for multi-index hashing. We derive novel training objectives, which enable to learn hash codes that reduce the candidate sets produced by multi-index hashing, while being end-to-end trainable. In fact, our proposed training objectives are model agnostic, i.e., not tied to how the hash codes are generated specifically in MISH, and are straight-forward to include in existing and future semantic hashing models. We experimentally compare MISH to state-of-the-art semantic hashing baselines in the task of document similarity search. We find that even though multi-index hashing also improves the efficiency of the baselines compared to a linear scan, they are still upwards of 33% slower than MISH, while MISH is still able to obtain state-of-the-art effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99994ea2-cf70-4a21-8d7f-885dc2dc8f8aCited by top-tier papers2
- Projected Hamming Dissimilarity for Bit-Level Importance Coding in Collaborative FilteringChristian Hansen, Casper Hansen, Jakob Grue Simonsen, Christina LiomaWWW 2021 · 10 citations
- Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic HashingLiyang He, Zhenya Huang, Jiayu Liu, Enhong Chen et al.WWW 2024 · 9 citations
Builds on2
- Content-aware Neural Hashing for Cold-start RecommendationCasper Hansen, Christian Hansen, Jakob Grue Simonsen, Stephen Alstrup et al.SIGIR 2020 · 29 citations
- Projected Hamming Dissimilarity for Bit-Level Importance Coding in Collaborative FilteringChristian Hansen, Casper Hansen, Jakob Grue Simonsen, Christina LiomaWWW 2021 · 10 citations
Related papers
- Statistical Model-driven Similarity Hashing: Bridging Modalities for Efficient Unsupervised RetrievalMingjin Kuai, Jun Long, Zhan YangAAAI 2025 · 3 citations
- Webly Supervised Image Hashing with Lightweight Semantic Transfer NetworkHui Cui, Lei Zhu, Jingjing Li, Zheng Zhang et al.ACM MM 2022 · 8 citations
- Deep Unsupervised Hybrid-similarity Hadamard HashingWanqian Zhang, Dayan Wu, Yu Zhou, Bo Li et al.ACM MM 2020 · 40 citations
- Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal DataXiao-Ming Wu, Xin Luo, Yu-Wei Zhan, Chenlu Ding et al.AAAI 2022 · 14 citations
- One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning ObjectiveJiun Tian Hoe, Kam Woh Ng, Tianyu Zhang, Chee Seng Chan et al.NeurIPS 2021 · 174 citations
