Efficient Implicit Unsupervised Text Hashing using Adversarial Autoencoder
Khoa D. Doan, Chandan K. Reddy
Abstract
Searching for documents with semantically similar content is a fundamental problem in the information retrieval domain with various challenges, primarily, in terms of efficiency and effectiveness. Despite the promise of modeling structured dependencies in documents, several existing text hashing methods lack an efficient mechanism to incorporate such vital information. Additionally, the desired characteristics of an ideal hash function, such as robustness to noise, low quantization error and bit balance/uncorrelation, are not effectively learned with existing methods. This is because of the requirement to either tune additional hyper-parameters or optimize these heuristically and explicitly constructed cost functions. In this paper, we propose a Denoising Adversarial Binary Autoencoder (DABA) model which presents a novel representation learning framework that captures structured representation of text documents in the learned hash function. Also, adversarial training provides an alternative direction to implicitly learn a hash function that captures all the desired characteristics of an ideal hash function. Essentially, DABA adopts a novel single-optimization adversarial training procedure that minimizes the Wasserstein distance in its primal domain to regularize the encoder’s output of either a recurrent neural network or a convolutional autoencoder. We empirically demonstrate the effectiveness of our proposed method in capturing the intrinsic semantic manifold of the related documents. The proposed method outperforms the current state-of-the-art shallow and deep unsupervised hashing methods for the document retrieval task on several prominent document collections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5912c96-8455-41ca-906c-beeb3a4cab68Cited by top-tier papers5
- LIRA: Learnable, Imperceptible and Robust Backdoor AttacksKhoa D. Doan, Yingjie Lao, Weijie Zhao, Ping LiICCV 2021 · 313 citations
- One Loss for Quantization: Deep Hashing with Discrete Wasserstein Distributional MatchingKhoa D. Doan, Peng Yang, Ping LiCVPR 2022 · 46 citations
- Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic HashingLiyang He, Zhenya Huang, Jiayu Liu, Enhong Chen et al.WWW 2024 · 9 citations
- Asymmetric Hashing for Fast Ranking via Neural Network MeasuresKhoa D. Doan, Shulong Tan, Weijie Zhao, Ping LiSIGIR 2023 · 3 citations
- Copyright-Certified Distillation Dataset: Distilling One Million Coins into One Bitcoin with Your Private KeyTengjun Liu, Ying Chen, Wanxuan GuAAAI 2023 · 1 citation
Related papers
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual InformationKarl Stratos, Sam WisemanICML 2020 · 9 citations
- Deep Unsupervised Hashing via External GuidanceQihong Song, XitingLiu, Hongyuan Zhu, Joey Tianyi Zhou et al.ICML 2025
- Auto-Encoding Twin-Bottleneck HashingYuming Shen, Jie Qin, Jiaxin Chen, Mengyang Yu et al.CVPR 2020
- Adaptive Structural Similarity Preserving for Unsupervised Cross Modal HashingLiang Li, Baihua Zheng, Weiwei SunACM MM 2022 · 29 citations
- Statistical Model-driven Similarity Hashing: Bridging Modalities for Efficient Unsupervised RetrievalMingjin Kuai, Jun Long, Zhan YangAAAI 2025 · 3 citations
