Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang
摘要
Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other tasks where supervision is limited, with the implicit assumption that it is possible to generalize from one task to all the rest. However, this overlooks the fact that there are many diverse and unique retrieval tasks, each targeting different search intents, queries, and search domains. In this paper, we suggest to work on Few-shot Dense Retrieval, a setting where each task comes with a short description and a few examples. To amplify the power of a few examples, we propose Promptbase Query Generation for Retriever (PROMPTAGATOR ), which leverages large language models (LLM) as a few-shot query generator, and creates task-specific retrievers based on the generated data. Powered by LLM's generalization ability, PROMPTAGATOR makes it possible to create task-specific end-to-end retrievers solely based on a few examples without using Natural Questions (Kwiatkowski et al., 2019) or MS MARCO (Nguyen et al., 2016) to train dual encoders. Surprisingly, LLM prompting with no more than 8 examples allows dual encoders to outperform heavily engineered models trained on MS MARCO like ColBERT v2 (Santhanam et al., 2022) by more than 1.2 nDCG on average on 11 retrieval sets. Further training standard-size re-rankers using the same generated data yields another 5.0 point nDCG improvement. Our studies determine that query generation can be far more effective than previously observed, especially when a small amount of task-specific knowledge is given.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper53
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang 等EMNLP 2023 · 被引用 182 次
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 等NeurIPS 2023 · 被引用 60 次
- Dense X Retrieval: What Retrieval Granularity Should We Use?Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu 等EMNLP 2024 · 被引用 52 次
- Ad Auctions for LLMs via Retrieval Augmented GenerationMohammadTaghi Hajiaghayi, Sébastien Lahaie, Keivan Rezaei, Suho ShinNeurIPS 2024 · 被引用 31 次
- Multiview Identifiers Enhanced Generative RetrievalYongqi Li, Nan Yang, Liang Wang, Furu Wei 等ACL 2023 · 被引用 30 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
相关 Paper
- From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query GenerationZhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao 等KDD 2025 · 被引用 4 次
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin 等EMNLP 2024 · 被引用 26 次
- HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval GeneralizationZefeng Cai, Chongyang Tao, Tao Shen, Can Xu 等ICLR 2023
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian 等EMNLP 2023 · 被引用 23 次
- Few-shot Reranking for Multi-hop QA via Language Model PromptingMuhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee 等ACL 2023 · 被引用 5 次
