UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts
Abstract
Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this challenge, we develop and motivate a method for using large language models (LLMs) to generate large numbers of synthetic queries cheaply. The method begins by generating a small number of synthetic queries using an expensive LLM. After that, a much less expensive one is used to create large numbers of synthetic queries, which are used to fine-tune a family of reranker models. These rerankers are then distilled into a single efficient retriever for use in the target domain. We show that this technique boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16ecd48b-2453-4391-bec1-9af33e2db154Cited by top-tier papers6
- Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented GenerationGuanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang et al.WWW 2025 · 44 citations
- Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERTJon Saad-Falcon, Daniel Y. Fu, Simran Arora, Neel Guha et al.ICML 2024 · 24 citations
- LargePiG for Hallucination-Free Query Generation: Your Large Language Model is Secretly a Pointer GeneratorZhongxiang Sun, Zihua Si, Xiaoxue Zang, Kai Zheng et al.WWW 2025 · 5 citations
- Incorporating Verification Standards for Security Requirements Generation from Functional SpecificationsXiaoli Lian, Shuaisong Wang, Hanyu Zou, Fang Liu et al.FSE 2025 · 4 citations
- On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme ClassificationJatin Prakash, Anirudh Buvanesh, Bishal Santra, Deepak Saini et al.KDD 2025 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Leveraging LLMs for Unsupervised Dense Retriever RankingEkaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, Guido ZucconSIGIR 2024 · 21 citations
- PANGEA: Projection-Based Augmentation with Non-Relevant General Data for Enhanced Domain Adaptation in LLMsSeungyoo Lee, Giung Nam, Moonseok Choi, Hyungi Lee et al.NeurIPS 2025
- Promptagator: Few-shot Dense Retrieval From 8 ExamplesZhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan et al.ICLR 2023 · 46 citations
- REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingPinhuan Wang, Zhiqiu Xia, Chunhua Liao, Feiyi Wang et al.EMNLP 2025
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
