Expand, Highlight, Generate: RL-driven Document Generation for Passage Reranking
Arian Askari, Mohammad Aliannejadi, Chuan Meng, Evangelos Kanoulas, Suzan Verberne
Abstract
<p>Generating synthetic training data based on large language models (LLMs) for ranking models has gained attention recently. Prior studies use LLMs to build pseudo query-document pairs by generating synthetic queries from documents in a corpus. In this paper, we propose a new perspective of data augmentation: generating synthetic documents from queries. To achieve this, we propose DocGen, that consists of a three-step pipeline that utilizes the few-shot capabilities of LLMs. DocGen pipeline performs synthetic document generation by (i) expanding, (ii) highlighting the original query, and then (iii) generating a synthetic document that is likely to be relevant to the query. To further improve the relevance between generated synthetic documents and their corresponding queries, we propose DocGen-RL, which regards the estimated relevance of the document as a reward and leverages reinforcement learning (RL) to optimize DocGen pipeline. Extensive experiments demonstrate that DocGen pipeline and DocGen-RL significantly outperform existing state-of-theart data augmentation methods, such as InPars, indicating that our new perspective of generating documents leverages the capacity of LLMs in generating synthetic data more effectively. We release the code, generated data, and model checkpoints to foster research in this area.<br></p>
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bcf5e6f-953a-42ce-afdd-95ae3dd34316Cited by top-tier papers1
Ask how each one uses itBuilds on11
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- Guiding Large Language Models via Directional Stimulus PromptingZekun Li, Baolin Peng, Pengcheng He, Michel Galley et al.NeurIPS 2023 · 163 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 88 citations
Related papers
- On Synthetic Data Strategies for Domain-Specific Generative RetrievalHaoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang et al.ACL 2025 · 6 citations
- q2d: Turning Questions into Dialogs to Teach Models How to SearchYonatan Bitton, Shlomi Cohen-Ganor, Ido Hakimi, Yoad Lewenberg et al.EMNLP 2023 · 3 citations
- DataGen: Unified Synthetic Dataset Generation via Large Language ModelsYue Huang, Siyuan Wu, Chujie Gao, Dongping Chen et al.ICLR 2025
- GENRA: Enhancing Zero-shot Retrieval with Rank AggregationGeorgios Katsimpras, Georgios PaliourasEMNLP 2024 · 1 citation
- LLM-powered Data Augmentation for Enhanced Cross-lingual PerformanceChenxi Whitehouse, Monojit Choudhury, Alham Fikri AjiEMNLP 2023 · 53 citations
