Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang
Abstract
Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other tasks where supervision is limited, with the implicit assumption that it is possible to generalize from one task to all the rest. However, this overlooks the fact that there are many diverse and unique retrieval tasks, each targeting different search intents, queries, and search domains. In this paper, we suggest to work on Few-shot Dense Retrieval, a setting where each task comes with a short description and a few examples. To amplify the power of a few examples, we propose Promptbase Query Generation for Retriever (PROMPTAGATOR ), which leverages large language models (LLM) as a few-shot query generator, and creates task-specific retrievers based on the generated data. Powered by LLM's generalization ability, PROMPTAGATOR makes it possible to create task-specific end-to-end retrievers solely based on a few examples without using Natural Questions (Kwiatkowski et al., 2019) or MS MARCO (Nguyen et al., 2016) to train dual encoders. Surprisingly, LLM prompting with no more than 8 examples allows dual encoders to outperform heavily engineered models trained on MS MARCO like ColBERT v2 (Santhanam et al., 2022) by more than 1.2 nDCG on average on 11 retrieval sets. Further training standard-size re-rankers using the same generated data yields another 5.0 point nDCG improvement. Our studies determine that query generation can be far more effective than previously observed, especially when a small amount of task-specific knowledge is given.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65629951-dee0-4254-a504-5ee43d0c3a18Cited by top-tier papers53
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei et al.NeurIPS 2023 · 60 citations
- Dense X Retrieval: What Retrieval Granularity Should We Use?Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu et al.EMNLP 2024 · 52 citations
- Ad Auctions for LLMs via Retrieval Augmented GenerationMohammadTaghi Hajiaghayi, Sébastien Lahaie, Keivan Rezaei, Suho ShinNeurIPS 2024 · 31 citations
- Multiview Identifiers Enhanced Generative RetrievalYongqi Li, Nan Yang, Liang Wang, Furu Wei et al.ACL 2023 · 30 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
Related papers
- From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query GenerationZhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao et al.KDD 2025 · 4 citations
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval GeneralizationZefeng Cai, Chongyang Tao, Tao Shen, Can Xu et al.ICLR 2023
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian et al.EMNLP 2023 · 23 citations
- Few-shot Reranking for Multi-hop QA via Language Model PromptingMuhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee et al.ACL 2023 · 5 citations
