Precise Zero-Shot Dense Retrieval without Relevance Labels
Luyu Gao, Xueguang Ma, Jimmy Lin, Jamie Callan
Abstract
While dense retrieval has been shown to be effective and efficient across tasks and languages, it remains difficult to create effective fully zero-shot dense retrieval systems when no relevance labels are available. In this paper, we recognize the difficulty of zero-shot learning and encoding relevance. Instead, we propose to pivot through Hypothetical Document Embeddings (HyDE). Given a query, HyDE first zero-shot prompts an instruction-following language model (e.g., InstructGPT) to generate a hypothetical document. The document captures relevance patterns but is "fake" and may contain hallucinations. Then, an unsupervised contrastively learned encoder (e.g., Contriever) encodes the document into an embedding vector. This vector identifies a neighborhood in the corpus embedding space, from which similar real documents are retrieved based on vector similarity. This second step grounds the generated document to the actual corpus, with the encoder's dense bottleneck filtering out the hallucinations. Our experiments show that HyDE significantly outperforms the state-ofthe-art unsupervised dense retriever Contriever and shows strong performance comparable to fine-tuned retrievers across various tasks (e.g. web search, QA, fact verification) and in non-English languages (e.g., sw, ko, ja, bn). 1 * Equal contribution. 1 No models were trained or fine-tuned in writing this paper. Our open-source code is available at https://github.com/ texttron/hyde .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers101
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval AugmentationHongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao et al.WWW 2025 · 92 citations
- Searching for Best Practices in Retrieval-Augmented GenerationXiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang et al.EMNLP 2024 · 75 citations
- RAGraph: A General Retrieval-Augmented Graph Learning FrameworkXinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang et al.NeurIPS 2024 · 42 citations
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang et al.EMNLP 2022 · 21 citations
- Graded Relevance Scoring of Written Essays with Dense RetrievalSalam Albatarni, Sohaila Eltanbouly, Tamer ElsayedSIGIR 2024 · 4 citations
- PESCO: Prompt-enhanced Self Contrastive Learning for Zero-shot Text ClassificationYau-Shian Wang, Ta-Chung Chi, Ruohong Zhang, Yiming YangACL 2023 · 17 citations
- Unsupervised Corpus Aware Language Model Pre-training for Dense Passage RetrievalLuyu Gao, Jamie CallanACL 2022
