From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query Generation
Zhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao, Xi Chen, Jingshuai Zhang, Chen Zhu, Hengshu Zhu
摘要
Document retrieval, designed to recall query-relevant documents from expansive collections, is essential for information-seeking tasks, such as web search and open-domain question-answering. Advances in representation learning and pretrained language models (PLMs) have driven a paradigm shift from traditional sparse retrieval methods to more effective dense retrieval approaches, forging enhanced semantic connections between queries and documents and establishing new performance benchmarks. However, reliance on extensive annotated document-query pairs limits their competitiveness in low-resource scenarios. Recent research efforts employing the few-shot capabilities of large language models (LLMs) and prompt engineering for synthetic data generation have emerged as a promising solution. Nonetheless, these approaches are hindered by the generation of lower-quality data within the conventional dense retrieval training process. To this end, in this paper, we introduce iGFT, a framework aimed at enhancing low-resource dense retrieval by integrating a three-phase process --- Generation, Filtering, and Tuning --- coupled with an iterative optimization strategy. Specifically, we first employ supervised fine-tuning on limited ground truth data, enabling an LLM to function as the generator capable of producing potential queries from given documents. Subsequently, we present a multi-stage filtering module to minimize noise in the generated data while retaining samples poised to significantly improve the dense retrieval model's performance in the follow-up fine-tuning process. Furthermore, we design a novel iterative optimization strategy that dynamically optimizes the query generator for producing more informative queries, thereby enhancing the efficacy of the entire framework. Finally, extensive experiments conducted on a series of publicly available retrieval benchmark datasets have demonstrated the effectiveness of the proposed iGFT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin 等NeurIPS 2025 · 被引用 7 次
- Beyond the Known: An Unknown-Aware Large Language Model for Open-Set Text ClassificationXi Chen, Chuan Qin, Ziqi Wang, Shasha Hu 等ICLR 2026
- TLSA: LLM-Guided Text-Label Space Alignment with Contrastive Learning for Generalized Category DiscoveryWenxi Xu, Chuan Qin, Xi Chen, Chuyu Fang 等ACL 2026
- GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category DiscoveryXi Chen, Chuan Qin, Jinpeng Li, Shasha Hu 等ACL 2026
相关 Paper
- Promptagator: Few-shot Dense Retrieval From 8 ExamplesZhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan 等ICLR 2023 · 被引用 46 次
- Synergistic Interplay between Search and Large Language Models for Information RetrievalJiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen 等ACL 2024 · 被引用 7 次
- MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question AnsweringXiusi Chen, Jyun-Yu Jiang, Wei-Cheng Chang, Cho-Jui Hsieh 等ACL 2024 · 被引用 7 次
- UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian 等EMNLP 2023 · 被引用 23 次
- RA-DIT: Retrieval-Augmented Dual Instruction TuningXi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi 等ICLR 2024 · 被引用 229 次
