Hybrid Inverted Index Is a Robust Accelerator for Dense Retrieval
Peitian Zhang, Zheng Liu, Shitao Xiao, Zhicheng Dou, Jing Yao
Abstract
Inverted file structure is a common technique for accelerating dense retrieval. It clusters documents based on their embeddings; during searching, it probes nearby clusters w.r.t. an input query and only evaluates documents within them by subsequent codecs, thus avoiding the expensive cost of exhaustive traversal. However, the clustering is always lossy, which results in the miss of relevant documents in the probed clusters and hence degrades retrieval quality. In contrast, lexical matching, such as overlaps of salient terms, tends to be strong feature for identifying relevant documents. In this work, we present the Hybrid Inverted Index (HI 2 ), where the embedding clusters and salient terms work collaboratively to accelerate dense retrieval. To make best of both effectiveness and efficiency, we devise a cluster selector and a term selector, to construct compact inverted lists and efficiently searching through them. Moreover, we leverage simple unsupervised algorithms as well as end-to-end knowledge distillation to learn these two modules, with the latter further boosting the effectiveness. Based on comprehensive experiments on popular retrieval benchmarks, we verify that clusters and terms indeed complement each other, enabling HI 2 to achieve lossless retrieval quality with competitive efficiency across various index settings. Our code and checkpoint are publicly available at https://github.com/ namespace-Pt/Adon/tree/HI2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generative Retrieval via Term Set GenerationPeitian Zhang, Zheng Liu, Yujia Zhou, Zhicheng Dou et al.SIGIR 2024 · 13 citations
- Threshold-driven Pruning with Segmented Maximum Term Weights for Approximate Cluster-based Sparse RetrievalYifan Qiao, Parker Carlson, Shanxiu He, Yingrui Yang et al.EMNLP 2024 · 7 citations
- FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor SearchZekai Wu, Jiabao Jin, Peng Cheng, Xiaoyao Zhong et al.SIGMOD 2026
- Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware SamplingYuanbang Liang, Jing Wu, Yu-Kun Lai, Yipeng QinICML 2024
Builds on9
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv et al.ICLR 2022 · 137 citations
- SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao et al.ACL 2023 · 41 citations
Related papers
- Lexically-Accelerated Dense RetrievalHrishikesh Kulkarni, Sean MacAvaney, Nazli Goharian, Ophir FriederSIGIR 2023 · 30 citations
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang et al.SIGIR 2022 · 31 citations
- Elastic Index Selection for Label-Hybrid AKNN SearchMingyu Yang, Wenxuan Xia, Wentao Li, Raymond Chi-Wing Wong et al.VLDB 2026
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang et al.ACL 2021
- LIDER: An Efficient High-dimensional Learned Index for Large-scale Dense Passage RetrievalYifan Wang, Haodi Ma, Daisy Zhe WangVLDB 2023 · 17 citations
