In defense of dual-encoders for neural ranking
Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim, Sashank J. Reddi, Sanjiv Kumar
摘要
Transformer-based models have proven successful in information retrieval problems, which seek to identify relevant documents for a given query. There are two broad flavours of such models: cross-attention (CA) models, which learn a joint query-document embedding, and dual-encoder (DE) models, which learn separate embeddings for the query and the document. Empirically, CA models are often more accurate, which has motivated several works that seek to bridge this performance gap. However, a fundamental question remains less explored: does this gap reflect a limitation in DE models' capacity, or training procedure? In this paper, we study this question, with three contributions. First, we establish theoretically that with a sufficiently large encoder size, DE models can capture a broad class of scores without cross-attention. Second, we show that on real-world problems, the gap between CA and DE models may be due to the latter overfitting to the training set. To mitigate this, we propose a distillation strategy that focuses on preserving the ordering amongst documents, and confirm its efficacy on neural re-ranking benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsYue Yu, Wei Ping, Zihan Liu, Boxin Wang 等NeurIPS 2024 · 被引用 321 次
- Fine-Grained Distillation for Long Document RetrievalYucheng Zhou, Tao Shen, Xiubo Geng, Chongyang Tao 等AAAI 2024 · 被引用 44 次
- Alternating Updates for Efficient TransformersCenk Baykal, Dylan J. Cutler, Nishanth Dikkala, Nikhil Ghosh 等NeurIPS 2023 · 被引用 13 次
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeChong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka 等NeurIPS 2025 · 被引用 4 次
- KwikBucks: Correlation Clustering with Cheap-Weak and Expensive-Strong SignalsSandeep Silwal, Sara Ahmadian, Andrew Nystrom, Andrew McCallum 等ICLR 2023 · 被引用 3 次
它引用的顶会 Paper13
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
相关 Paper
- USTAD: Unified Single-model Training Achieving Diverse Scores for Information RetrievalSeungyeon Kim, Ankit Singh Rawat, Manzil Zaheer, Wittawat Jitkrittum 等ICML 2024 · 被引用 3 次
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval With TransformersAntoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic 等CVPR 2021
- How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?Yuxin Chen, Zongyang Ma, Ziqi Zhang, Zhongang Qi 等CVPR 2024
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong 等EMNLP 2022 · 被引用 6 次
- Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-EncodersNishant Yadav, Nicholas Monath, Manzil Zaheer, Rob Fergus 等ICLR 2024 · 被引用 2 次
