Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization
Zexuan Qiu, Qinliang Su, Jianxing Yu, Shijing Si
Abstract
Efficient document retrieval heavily relies on the technique of semantic hashing, which learns a binary code for every document and employs Hamming distance to evaluate document distances. However, existing semantic hashing methods are mostly established on outdated TFIDF features, which obviously do not contain lots of important semantic information about documents. Furthermore, the Hamming distance can only be equal to one of several integer values, significantly limiting its representational ability for document distances. To address these issues, in this paper, we propose to leverage BERT embeddings to perform efficient retrieval based on the product quantization technique, which will assign for every document a real-valued codeword from the codebook, instead of a binary code as in semantic hashing. Specifically, we first transform the original BERT embeddings via a learnable mapping and feed the transformed embedding into a probabilistic product quantization module to output the assigned codeword. The refining and quantizing modules can be optimized in an end-to-end manner by minimizing the probabilistic contrastive loss. A mutual information maximization based method is further proposed to improve the representativeness of codewords, so that documents can be quantized more accurately. Extensive experiments conducted on three benchmarks demonstrate that our proposed method significantly outperforms current state-of-the-art baselines 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e15b213-6b8d-4fd4-a0e7-acc0a0a64adeCited by top-tier papers1
Ask how each one uses itBuilds on9
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Self-supervised Product Quantization for Deep Unsupervised Image RetrievalYoung Kyun Jang, Nam Ik ChoICCV 2021 · 90 citations
- Contrastive Quantization with Code Memory for Unsupervised Image RetrievalJinpeng Wang, Ziyun Zeng, Bin Chen, Tao Dai et al.AAAI 2022 · 56 citations
Related papers
- Unsupervised Multi-Index Semantic HashingChristian Hansen, Casper Hansen, Jakob Grue Simonsen, Stephen Alstrup et al.WWW 2021 · 11 citations
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang et al.NeurIPS 2023 · 151 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Distilling Large Embeddings via Hyperspherical Householder QuantizationYihang Wang, Bin Wu, Yueyang Su, Tianfu Zhang et al.ACL 2026
- One Loss for Quantization: Deep Hashing with Discrete Wasserstein Distributional MatchingKhoa D. Doan, Peng Yang, Ping LiCVPR 2022 · 46 citations
