From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query Generation
Zhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao, Xi Chen, Jingshuai Zhang, Chen Zhu, Hengshu Zhu
Abstract
Document retrieval, designed to recall query-relevant documents from expansive collections, is essential for information-seeking tasks, such as web search and open-domain question-answering. Advances in representation learning and pretrained language models (PLMs) have driven a paradigm shift from traditional sparse retrieval methods to more effective dense retrieval approaches, forging enhanced semantic connections between queries and documents and establishing new performance benchmarks. However, reliance on extensive annotated document-query pairs limits their competitiveness in low-resource scenarios. Recent research efforts employing the few-shot capabilities of large language models (LLMs) and prompt engineering for synthetic data generation have emerged as a promising solution. Nonetheless, these approaches are hindered by the generation of lower-quality data within the conventional dense retrieval training process. To this end, in this paper, we introduce iGFT, a framework aimed at enhancing low-resource dense retrieval by integrating a three-phase process --- Generation, Filtering, and Tuning --- coupled with an iterative optimization strategy. Specifically, we first employ supervised fine-tuning on limited ground truth data, enabling an LLM to function as the generator capable of producing potential queries from given documents. Subsequently, we present a multi-stage filtering module to minimize noise in the generated data while retaining samples poised to significantly improve the dense retrieval model's performance in the follow-up fine-tuning process. Furthermore, we design a novel iterative optimization strategy that dynamically optimizes the query generator for producing more informative queries, thereby enhancing the efficacy of the entire framework. Finally, extensive experiments conducted on a series of publicly available retrieval benchmark datasets have demonstrated the effectiveness of the proposed iGFT.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 05770dbb-ae13-45e9-87f0-dbbc394239f9Cited by top-tier papers4
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin et al.NeurIPS 2025 · 7 citations
- Beyond the Known: An Unknown-Aware Large Language Model for Open-Set Text ClassificationXi Chen, Chuan Qin, Ziqi Wang, Shasha Hu et al.ICLR 2026
- TLSA: LLM-Guided Text-Label Space Alignment with Contrastive Learning for Generalized Category DiscoveryWenxi Xu, Chuan Qin, Xi Chen, Chuyu Fang et al.ACL 2026
- GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category DiscoveryXi Chen, Chuan Qin, Jinpeng Li, Shasha Hu et al.ACL 2026
Related papers
- Promptagator: Few-shot Dense Retrieval From 8 ExamplesZhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan et al.ICLR 2023 · 46 citations
- Synergistic Interplay between Search and Large Language Models for Information RetrievalJiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen et al.ACL 2024 · 7 citations
- MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question AnsweringXiusi Chen, Jyun-Yu Jiang, Wei-Cheng Chang, Cho-Jui Hsieh et al.ACL 2024 · 7 citations
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian et al.EMNLP 2023 · 23 citations
- RA-DIT: Retrieval-Augmented Dual Instruction TuningXi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi et al.ICLR 2024 · 229 citations
