Contextual Document Embeddings
John Xavier Morris, Alexander M. Rush
Abstract
Dense document embeddings are central to neural retrieval. The dominant paradigm is to train and construct embeddings by running encoders directly on individual documents. In this work, we argue that these embeddings, while effective, are implicitly out-of-context for targeted use cases of retrieval, and that a document embedding should take into account both the document and neighboring documents in context -analogous to contextualized word embeddings. We propose two complementary methods for contextualized document embeddings: first, an alternative contrastive learning objective that explicitly incorporates document neighbors into the intra-batch contextual loss; second, a new contextual architecture that explicitly encodes neighbor document information into the encoded representation. Results show that both methods achieve better performance than biencoders in several settings, with differences especially pronounced out-of-domain. We achieve stateof-the-art results on the MTEB benchmark with no hard negative mining, score distillation, dataset-specific instructions, intra-GPU example-sharing, or extremely large batch sizes. Our method can be applied to improve performance on any contrastive learning dataset and any biencoder.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K et al.NeurIPS 2025 · 30 citations
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversXueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin et al.ACL 2025 · 20 citations
- Training compute-optimal transformer encoder modelsMegi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCunEMNLP 2025 · 1 citation
- LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense RetrievalYanzhen Shen, Sihao Chen, Xueqiang Xu, Yunyi Zhang et al.EMNLP 2025 · 1 citation
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsMax Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut et al.EMNLP 2025 · 1 citation
Builds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
Related papers
- CODER: An efficient framework for improving retrieval through COntextual Document Embedding RerankingGeorge Zerveas, Navid Rekabsaz, Daniel Cohen, Carsten EickhoffEMNLP 2022 · 10 citations
- Leveraging Cognitive Complexity of Texts for Contextualization in Dense RetrievalEffrosyni Sokli, Georgios Peikos, Pranav Kasela, Gabriella PasiEMNLP 2025
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding ModelsChankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman et al.ICLR 2025
- NewsEmbed: Modeling News through Pre-trained Document RepresentationsJialu Liu, Tianqi Liu, Cong YuKDD 2021 · 15 citations
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan et al.EMNLP 2025 · 2 citations
