Hybrid Inverted Index Is a Robust Accelerator for Dense Retrieval
Peitian Zhang, Zheng Liu, Shitao Xiao, Zhicheng Dou, Jing Yao
摘要
Inverted file structure is a common technique for accelerating dense retrieval. It clusters documents based on their embeddings; during searching, it probes nearby clusters w.r.t. an input query and only evaluates documents within them by subsequent codecs, thus avoiding the expensive cost of exhaustive traversal. However, the clustering is always lossy, which results in the miss of relevant documents in the probed clusters and hence degrades retrieval quality. In contrast, lexical matching, such as overlaps of salient terms, tends to be strong feature for identifying relevant documents. In this work, we present the Hybrid Inverted Index (HI 2 ), where the embedding clusters and salient terms work collaboratively to accelerate dense retrieval. To make best of both effectiveness and efficiency, we devise a cluster selector and a term selector, to construct compact inverted lists and efficiently searching through them. Moreover, we leverage simple unsupervised algorithms as well as end-to-end knowledge distillation to learn these two modules, with the latter further boosting the effectiveness. Based on comprehensive experiments on popular retrieval benchmarks, we verify that clusters and terms indeed complement each other, enabling HI 2 to achieve lossless retrieval quality with competitive efficiency across various index settings. Our code and checkpoint are publicly available at https://github.com/ namespace-Pt/Adon/tree/HI2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generative Retrieval via Term Set GenerationPeitian Zhang, Zheng Liu, Yujia Zhou, Zhicheng Dou 等SIGIR 2024 · 被引用 13 次
- Threshold-driven Pruning with Segmented Maximum Term Weights for Approximate Cluster-based Sparse RetrievalYifan Qiao, Parker Carlson, Shanxiu He, Yingrui Yang 等EMNLP 2024 · 被引用 7 次
- FGIM: a Fast Graph-based Indexes Merging Framework for Approximate Nearest Neighbor SearchZekai Wu, Jiabao Jin, Peng Cheng, Xiaoyao Zhong 等SIGMOD 2026
- Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware SamplingYuanbang Liang, Jing Wu, Yu-Kun Lai, Yipeng QinICML 2024
它引用的顶会 Paper9
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv 等ICLR 2022 · 被引用 137 次
- SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao 等ACL 2023 · 被引用 41 次
相关 Paper
- Lexically-Accelerated Dense RetrievalHrishikesh Kulkarni, Sean MacAvaney, Nazli Goharian, Ophir FriederSIGIR 2023 · 被引用 30 次
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang 等SIGIR 2022 · 被引用 31 次
- Elastic Index Selection for Label-Hybrid AKNN SearchMingyu Yang, Wenxuan Xia, Wentao Li, Raymond Chi-Wing Wong 等VLDB 2026
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang 等ACL 2021
- LIDER: An Efficient High-dimensional Learned Index for Large-scale Dense Passage RetrievalYifan Wang, Haodi Ma, Daisy Zhe WangVLDB 2023 · 被引用 17 次
