Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix Factorization
Nishant Yadav, Nicholas Monath, Rico Angell, Manzil Zaheer, Andrew McCallum
Abstract
Efficient k-nearest neighbor search is a fundamental task, foundational for many problems in NLP. When the similarity is measured by dot-product between dual-encoder vectors or ℓ 2 -distance, there already exist many scalable and efficient search methods. But not so when similarity is measured by more accurate and expensive black-box neural similarity models, such as cross-encoders, which jointly encode the query and candidate neighbor. The cross-encoders' high computational cost typically limits their use to reranking candidates retrieved by a cheaper model, such as dual encoder or TF-IDF. However, the accuracy of such a two-stage approach is upper-bounded by the recall of the initial candidate set, and potentially requires additional training to align the auxiliary retrieval model with the cross-encoder model. In this paper, we present an approach that avoids the use of a dual-encoder for retrieval, relying solely on the cross-encoder. Retrieval is made efficient with CUR decomposition, a matrix decomposition approach that approximates all pairwise cross-encoder distances from a small subset of rows and columns of the distance matrix. Indexing items using our approach is computationally cheaper than training an auxiliary dual-encoder model through distillation. Empirically, for k > 10, our approach provides test-time recall-vs-computational cost trade-offs superior to the current widely-used methods that re-rank items retrieved using a dual-encoder or TF-IDF. Score Query Embedding Item Embedding Score Joint Query-Item Embedding Linear Layer Score Contextualized Query Embedding Contextualized Item Embedding Query (q) … … … Item (i) … … … Special Tokens Query & Item Tokens a) Dual-Encoder b) [CLS]-Cross-Encoder c) [EMB]-Cross-Encoder Model (e.g. BERT) Model (e.g. BERT) Model (e.g. BERT) Model (e.g. BERT) (a) Model architecture -10 0 10 Query-Item Score 10 -3 10 -2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam GenerationGauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent CallotICML 2024 · 35 citations
- A Bi-metric Framework for Efficient Nearest Neighbor SearchHaike Xu, Sandeep Silwal, Piotr IndykICML 2026 · 3 citations
- Comparing Neighbors Together Makes it Easy: Jointly Comparing Multiple Candidates for Efficient and Effective RetrievalJonghyun Song, Cheyon Jin, Wenlong Zhao, Andrew McCallum et al.EMNLP 2024 · 3 citations
- USTAD: Unified Single-model Training Achieving Diverse Scores for Information RetrievalSeungyeon Kim, Ankit Singh Rawat, Manzil Zaheer, Wittawat Jitkrittum et al.ICML 2024 · 3 citations
- Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-EncodersNishant Yadav, Nicholas Monath, Manzil Zaheer, Rob Fergus et al.ICLR 2024 · 2 citations
Builds on8
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillationsFangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz et al.ICLR 2022 · 36 citations
Related papers
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic et al.CVPR 2021
- Efficient Re-ranking with Cross-encoders via Early ExitFrancesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando et al.SIGIR 2025 · 9 citations
- Improving the Accuracy of Dense Retrieval on the Quantized Indexes via Gradient Optimization of the Target EmbeddingsCong Tan, Yongqi Shao, Hong Huo, Tao FangAAAI 2026
- Fast Deterministic CUR Matrix Decomposition with Accuracy AssuranceYasutoshi Ida, Sekitoshi Kanai, Yasuhiro Fujiwara, Tomoharu Iwata et al.ICML 2020 · 14 citations
- In defense of dual-encoders for neural rankingAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim et al.ICML 2022 · 29 citations
