Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings
Max Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut, Céline Hudelot, Pierre Colombo
摘要
A limitation of modern document retrieval embedding methods is that they typically encode passages (chunks) from the same documents independently, often overlooking crucial contextual information from the rest of the document that could greatly improve individual chunk representations. In this work, we introduce ConTEB (Contextaware Text Embedding Benchmark), a benchmark designed to evaluate retrieval models on their ability to leverage document-wide context. Our results show that state-of-the-art embedding models struggle in retrieval scenarios where context is required. To address this limitation, we propose InSeNT (In-sequence Negative Training), a novel contrastive posttraining approach which combined with late chunking pooling enhances contextual representation learning while preserving computational efficiency. Our method significantly improves retrieval quality on ConTEB without sacrificing base model performance. We further find chunks embedded with our method are more robust to suboptimal chunking strategies and larger retrieval corpus sizes. We opensource all artifacts at https://github.com/ illuin-tech/contextual-embeddings .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison 等ICML 2026 · 被引用 17 次
- ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World ScenariosAntónio Loison, Quentin Macé, Antoine Edy, Victor Xing 等ACL 2026 · 被引用 15 次
- Embedding-Based Context-Aware RerankerYe Yuan, Amin Shabani, Siqi LiuICLR 2026
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna 等ICLR 2024 · 被引用 460 次
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 被引用 187 次
相关 Paper
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
- Sentence-aware Contrastive Learning for Open-Domain Passage RetrievalWu Hong, Zhuosheng Zhang, Jinyuan Wang, Hai ZhaoACL 2022
- DAPR: A Benchmark on Document-Aware Passage RetrievalKexin Wang, Nils Reimers, Iryna GurevychACL 2024
- CODER: An efficient framework for improving retrieval through COntextual Document Embedding RerankingGeorge Zerveas, Navid Rekabsaz, Daniel Cohen, Carsten EickhoffEMNLP 2022 · 被引用 10 次
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin 等EMNLP 2023 · 被引用 2 次
