LIDER: An Efficient High-dimensional Learned Index for Large-scale Dense Passage Retrieval
Yifan Wang, Haodi Ma, Daisy Zhe Wang
Abstract
Passage retrieval has been studied for decades, and many recent approaches of passage retrieval are using dense embeddings generated from deep neural models, called "dense passage retrieval". The state-of-the-art end-to-end dense passage retrieval systems normally deploy a deep neural model followed by an approximate nearest neighbor (ANN) search module. The model generates embeddings of the corpus and queries, which are then indexed and searched by the high-performance ANN module. With the increasing data scale, the ANN module unavoidably becomes the bottleneck on efficiency. An alternative is the learned index, which achieves significantly high search efficiency by learning the data distribution and predicting the target data location. But most of the existing learned indexes are designed for low dimensional data, which are not suitable for dense passage retrieval with high-dimensional dense embeddings.
In this paper, we propose LIDER , an efficient high-dimensional L earned I ndex for large-scale DE nse passage R etrieval. LIDER has a clustering-based hierarchical architecture formed by two layers of core models. As the basic unit of LIDER to index and search data, a core model includes an adapted recursive model index (RMI) and a dimension reduction component which consists of an extended SortingKeys-LSH (SK-LSH) and a key re-scaling module. The dimension reduction component reduces the high-dimensional dense embeddings into one-dimensional keys and sorts them in a specific order, which are then used by the RMI to make fast prediction. Experiments show that LIDER has a higher search speed with high retrieval quality comparing to the state-of-the-art ANN indexes on passage retrieval tasks, e.g., on large-scale data it achieves 1.2x search speed and significantly higher retrieval quality than the fastest baseline in our evaluation. Furthermore, LIDER has a better capability of speed-quality trade-off.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor SearchJianyang Gao, Cheng LongSIGMOD 2024 · 83 citations
- Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High-Dimensional Vector Similarity Search on Data SegmentMengzhao Wang, Weizhi Xu, Xiaomeng Yi, Songlin Wu et al.SIGMOD 2024 · 63 citations
- A Survey for Efficient Open Domain Question AnsweringQin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao et al.ACL 2023 · 28 citations
- Accelerating String-key Learned Index Structures via Memoization-based Incremental TrainingMinsu Kim, Jinwoo Hwang, Guseul Heo, Seiyeon Cho et al.VLDB 2024 · 10 citations
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid SearchMengzhao Wang, Boyu Tan, Yunjun Gao, Hai Jin et al.VLDB 2026 · 8 citations
Builds on8
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 362 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
Related papers
- Lexically-Accelerated Dense RetrievalHrishikesh Kulkarni, Sean MacAvaney, Nazli Goharian, Ophir FriederSIGIR 2023 · 30 citations
- Leveraging Passage Embeddings for Efficient Listwise Reranking with Large Language ModelsQi Liu, Bo Wang, Nan Wang, Jiaxin MaoWWW 2025 · 26 citations
- Efficiently Learning Spatial IndicesGuanli Liu, Jianzhong Qi, Christian S. Jensen, James Bailey et al.ICDE 2023 · 12 citations
- SOLAR: Sparse Orthogonal Learned and Random EmbeddingsTharun Medini, Beidi Chen, Anshumali ShrivastavaICLR 2021 · 10 citations
- HIRE: A Hybrid Learned Index for Robust and Efficient Performance under Mixed WorkloadsXinyi Zhang, Liang Liang, Anastasia Ailamaki, Jianliang XuSIGMOD 2026 · 2 citations
