DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine
Yifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen, Qiaoqiao She, Jing Liu, Hua Wu, Haifeng Wang
Abstract
In this paper, we present DuReader retrieval , a large-scale Chinese dataset for passage retrieval. DuReader retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine. To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, we (1) reduce the false negatives in development and test sets by manually annotating results pooled from multiple retrievers, and (2) remove the training queries that are semantically similar to the development and testing queries. Additionally, we provide two outof-domain testing sets for cross-domain evaluation, as well as a set of human translated queries for for cross-lingual retrieval evaluation. The experiments demonstrate that DuReader retrieval is challenging and a number of problems remain unsolved, such as the salient phrase mismatch and the syntactic mismatch between queries and paragraphs. These experiments also show that dense retrievers do not generalize well across domains, and cross-lingual retrieval is essentially challenging. DuReader retrieval is publicly available at https://github.com/baidu/DuReader/ tree/master/DuReader-Retrieval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text EmbeddingsTengyu Pan, Zhichao Duan, Zhenyu Li, Bowen Dong et al.ACL 2025 · 3 citations
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query InferenceGuangyuan Ma, Yongliang Ma, Xuanrui Gou, Zhenpeng Su et al.ICLR 2026 · 3 citations
- Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate PairsCheng Gao, Chaojun Xiao, Zhenghao Liu, Huimin Chen et al.EMNLP 2024 · 1 citation
- Generative Representational Instruction TuningNiklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang et al.ICLR 2025
- Making Text Embedders Few-Shot LearnersChaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen et al.ICLR 2025
Builds on11
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- DAPR: A Benchmark on Document-Aware Passage RetrievalKexin Wang, Nils Reimers, Iryna GurevychACL 2024
- Recovering Gold from Black Sand: Multilingual Dense Passage Retrieval with Hard and False Negative SamplesTianhao Shen, Mingtong Liu, Ming Zhou, Deyi XiongEMNLP 2022 · 3 citations
- DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL DatasetLijie Wang, Ao Zhang, Kun Wu, Ke Sun et al.EMNLP 2020 · 40 citations
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan et al.EMNLP 2025 · 2 citations
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
