In defense of dual-encoders for neural ranking
Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim, Sashank J. Reddi, Sanjiv Kumar
Abstract
Transformer-based models have proven successful in information retrieval problems, which seek to identify relevant documents for a given query. There are two broad flavours of such models: cross-attention (CA) models, which learn a joint query-document embedding, and dual-encoder (DE) models, which learn separate embeddings for the query and the document. Empirically, CA models are often more accurate, which has motivated several works that seek to bridge this performance gap. However, a fundamental question remains less explored: does this gap reflect a limitation in DE models' capacity, or training procedure? In this paper, we study this question, with three contributions. First, we establish theoretically that with a sufficiently large encoder size, DE models can capture a broad class of scores without cross-attention. Second, we show that on real-world problems, the gap between CA and DE models may be due to the latter overfitting to the training set. To mitigate this, we propose a distillation strategy that focuses on preserving the ordering amongst documents, and confirm its efficacy on neural re-ranking benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 859b3ee3-cb69-4ff2-a74f-e9eb941a3a01Cited by top-tier papers12
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsYue Yu, Wei Ping, Zihan Liu, Boxin Wang et al.NeurIPS 2024 · 321 citations
- Fine-Grained Distillation for Long Document RetrievalYucheng Zhou, Tao Shen, Xiubo Geng, Chongyang Tao et al.AAAI 2024 · 44 citations
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh et al.NeurIPS 2023 · 13 citations
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeChong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka et al.NeurIPS 2025 · 4 citations
- KwikBucks: Correlation Clustering with Cheap-Weak and Expensive-Strong SignalsSandeep Silwal, Sara Ahmadian, Andrew Nystrom, Andrew McCallum et al.ICLR 2023 · 3 citations
Builds on13
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
Related papers
- USTAD: Unified Single-model Training Achieving Diverse Scores for Information RetrievalSeungyeon Kim, Ankit Singh Rawat, Manzil Zaheer, Wittawat Jitkrittum et al.ICML 2024 · 3 citations
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic et al.CVPR 2021
- How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?Yuxin Chen, Zongyang Ma, Ziqi Zhang, Zhongang Qi et al.CVPR 2024
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong et al.EMNLP 2022 · 6 citations
- Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-EncodersNishant Yadav, Nicholas Monath, Manzil Zaheer, Rob Fergus et al.ICLR 2024 · 2 citations
