DAPR: A Benchmark on Document-Aware Passage Retrieval
Kexin Wang, Nils Reimers, Iryna Gurevych
Abstract
The work of neural retrieval so far focuses on ranking short texts and is challenged with long documents. There are many cases where the users want to find a relevant passage within a long document from a huge corpus, e.g. Wikipedia articles, research papers, etc. We propose and name this task Document-Aware Passage Retrieval (DAPR). While analyzing the errors of the State-of-The-Art (SoTA) passage retrievers, we find the major errors (53.5%) are due to missing document context. This drives us to build a benchmark for this task including multiple datasets from heterogeneous domains. In the experiments, we extend the SoTA passage retrievers with document context via (1) hybrid retrieval with BM25 and (2) contextualized passage representations, which inform the passage representation with document context. We find despite that hybrid retrieval performs the strongest on the mixture of the easy and the hard queries, it completely fails on the hard queries that require document-context understanding. On the other hand, contextualized passage representations (e.g. prepending document titles) achieve good improvement on these hard queries, but overall they also perform rather poorly. Our created benchmark enables future research on developing and comparing retrieval systems for the new task. The code and the data are available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82129249-780a-4e9a-8451-73362bdbfcadCited by top-tier papers3
- Mixture-of-RAG: Integrating Text and Tables with Large Language ModelsChi Zhang, Qiyang Chen, Mengqi ZhangKDD 2026 · 1 citation
- SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization FrameworkCheng Guo, Hu Kai, Shuxian Liang, Yiyang Jiang et al.ACL 2025
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora et al.ACL 2025
Builds on8
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 63 citations
- ConditionalQA: A Complex Reading Comprehension Dataset with Conditional AnswersHaitian Sun, William W. Cohen, Ruslan SalakhutdinovACL 2022 · 41 citations
Related papers
- DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineYifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen et al.EMNLP 2022 · 10 citations
- Cross-document Event Coreference Search: Task, Dataset and ModelingAlon Eirew, Avi Caciularu, Ido DaganEMNLP 2022 · 3 citations
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsMax Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut et al.EMNLP 2025 · 1 citation
- Multi-Task Retrieval for Knowledge-Intensive TasksJean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen-tau Yih et al.ACL 2021
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
