Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix Factorization
Nishant Yadav, Nicholas Monath, Rico Angell, Manzil Zaheer, Andrew McCallum
摘要
Efficient k-nearest neighbor search is a fundamental task, foundational for many problems in NLP. When the similarity is measured by dot-product between dual-encoder vectors or ℓ 2 -distance, there already exist many scalable and efficient search methods. But not so when similarity is measured by more accurate and expensive black-box neural similarity models, such as cross-encoders, which jointly encode the query and candidate neighbor. The cross-encoders' high computational cost typically limits their use to reranking candidates retrieved by a cheaper model, such as dual encoder or TF-IDF. However, the accuracy of such a two-stage approach is upper-bounded by the recall of the initial candidate set, and potentially requires additional training to align the auxiliary retrieval model with the cross-encoder model. In this paper, we present an approach that avoids the use of a dual-encoder for retrieval, relying solely on the cross-encoder. Retrieval is made efficient with CUR decomposition, a matrix decomposition approach that approximates all pairwise cross-encoder distances from a small subset of rows and columns of the distance matrix. Indexing items using our approach is computationally cheaper than training an auxiliary dual-encoder model through distillation. Empirically, for k > 10, our approach provides test-time recall-vs-computational cost trade-offs superior to the current widely-used methods that re-rank items retrieved using a dual-encoder or TF-IDF. Score Query Embedding Item Embedding Score Joint Query-Item Embedding Linear Layer Score Contextualized Query Embedding Contextualized Item Embedding Query (q) … … … Item (i) … … … Special Tokens Query & Item Tokens a) Dual-Encoder b) [CLS]-Cross-Encoder c) [EMB]-Cross-Encoder Model (e.g. BERT) Model (e.g. BERT) Model (e.g. BERT) Model (e.g. BERT) (a) Model architecture -10 0 10 Query-Item Score 10 -3 10 -2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam GenerationGauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent CallotICML 2024 · 被引用 35 次
- A Bi-metric Framework for Efficient Nearest Neighbor SearchHaike Xu, Sandeep Silwal, Piotr IndykICML 2026 · 被引用 3 次
- Comparing Neighbors Together Makes it Easy: Jointly Comparing Multiple Candidates for Efficient and Effective RetrievalJonghyun Song, Cheyon Jin, Wenlong Zhao, Andrew McCallum 等EMNLP 2024 · 被引用 3 次
- USTAD: Unified Single-model Training Achieving Diverse Scores for Information RetrievalSeungyeon Kim, Ankit Singh Rawat, Manzil Zaheer, Wittawat Jitkrittum 等ICML 2024 · 被引用 3 次
- Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-EncodersNishant Yadav, Nicholas Monath, Manzil Zaheer, Rob Fergus 等ICLR 2024 · 被引用 2 次
它引用的顶会 Paper8
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel 等EMNLP 2020 · 被引用 336 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillationsFangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz 等ICLR 2022 · 被引用 36 次
相关 Paper
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic 等CVPR 2021
- Efficient Re-ranking with Cross-encoders via Early ExitFrancesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando 等SIGIR 2025 · 被引用 9 次
- Improving the Accuracy of Dense Retrieval on the Quantized Indexes via Gradient Optimization of the Target EmbeddingsCong Tan, Yongqi Shao, Hong Huo, Tao FangAAAI 2026
- Fast Deterministic CUR Matrix Decomposition with Accuracy AssuranceYasutoshi Ida, Sekitoshi Kanai, Yasuhiro Fujiwara, Tomoharu Iwata 等ICML 2020 · 被引用 14 次
- In defense of dual-encoders for neural rankingAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim 等ICML 2022 · 被引用 29 次
