CODER: An efficient framework for improving retrieval through COntextual Document Embedding Reranking
George Zerveas, Navid Rekabsaz, Daniel Cohen, Carsten Eickhoff
Abstract
Contrastive learning has been the dominant approach to training dense retrieval models. In this work, we investigate the impact of ranking context - an often overlooked aspect of learning dense retrieval models. In particular, we examine the effect of its constituent parts: jointly scoring a large number of negatives per query, using retrieved (query-specific) instead of random negatives, and a fully list-wise loss.To incorporate these factors into training, we introduce Contextual Document Embedding Reranking (CODER), a highly efficient retrieval framework. When reranking, it incurs only a negligible computational overhead on top of a first-stage method at run time (approx. 5 ms delay per query), allowing it to be easily combined with any state-of-the-art dual encoder method. Models trained through CODER can also be used as stand-alone retrievers.Evaluating CODER in a large set of experiments on the MS MARCO and TripClick collections, we show that the contextual reranking of precomputed document embeddings leads to a significant improvement in retrieval performance. This improvement becomes even more pronounced when more relevance information per query is available, shown in the TripClick collection, where we establish new state-of-the-art results by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bd4ae5d-6d4d-4894-9b7d-5b279f18c9f9Cited by top-tier papers2
- Enhancing the Ranking Context of Dense Retrieval through Reciprocal Nearest NeighborsGeorge Zerveas, Navid Rekabsaz, Carsten EickhoffEMNLP 2023 · 4 citations
- Leveraging Cognitive Complexity of Texts for Contextualization in Dense RetrievalEffrosyni Sokli, Georgios Peikos, Pranav Kasela, Gabriella PasiEMNLP 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
Related papers
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv et al.ICLR 2022 · 137 citations
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsMax Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut et al.EMNLP 2025 · 1 citation
- Recasting Web-Scale Query Suggestion as dense retrieval: Efficient, Up-to-Date, and Context-Aware SuggestionsSosuke Nishikawa, Naoki Yoshinaga, Nobuhiro KajiSIGIR 2026
- DVCQR: Dual-View Conversational Query Rewriting with Stage-wise Reinforcement LearningChenyi Li, Xinhui Tu, Zaixiang WangACL 2026
