DAPR: A Benchmark on Document-Aware Passage Retrieval
Kexin Wang, Nils Reimers, Iryna Gurevych
摘要
The work of neural retrieval so far focuses on ranking short texts and is challenged with long documents. There are many cases where the users want to find a relevant passage within a long document from a huge corpus, e.g. Wikipedia articles, research papers, etc. We propose and name this task Document-Aware Passage Retrieval (DAPR). While analyzing the errors of the State-of-The-Art (SoTA) passage retrievers, we find the major errors (53.5%) are due to missing document context. This drives us to build a benchmark for this task including multiple datasets from heterogeneous domains. In the experiments, we extend the SoTA passage retrievers with document context via (1) hybrid retrieval with BM25 and (2) contextualized passage representations, which inform the passage representation with document context. We find despite that hybrid retrieval performs the strongest on the mixture of the easy and the hard queries, it completely fails on the hard queries that require document-context understanding. On the other hand, contextualized passage representations (e.g. prepending document titles) achieve good improvement on these hard queries, but overall they also perform rather poorly. Our created benchmark enables future research on developing and comparing retrieval systems for the new task. The code and the data are available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mixture-of-RAG: Integrating Text and Tables with Large Language ModelsChi Zhang, Qiyang Chen, Mengqi ZhangKDD 2026 · 被引用 1 次
- SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization FrameworkCheng Guo, Hu Kai, Shuxian Liang, Yiyang Jiang 等ACL 2025
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora 等ACL 2025
它引用的顶会 Paper8
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 被引用 316 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 被引用 63 次
- ConditionalQA: A Complex Reading Comprehension Dataset with Conditional AnswersHaitian Sun, William W. Cohen, Ruslan SalakhutdinovACL 2022 · 被引用 41 次
相关 Paper
- DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineYifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen 等EMNLP 2022 · 被引用 10 次
- Cross-document Event Coreference Search: Task, Dataset and ModelingAlon Eirew, Avi Caciularu, Ido DaganEMNLP 2022 · 被引用 3 次
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsMax Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut 等EMNLP 2025 · 被引用 1 次
- Multi-Task Retrieval for Knowledge-Intensive TasksJean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen-tau Yih 等ACL 2021
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin 等EMNLP 2023 · 被引用 2 次
