On Synthetic Data Strategies for Domain-Specific Generative Retrieval
Haoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang, Zhiguo Wang
Abstract
This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the first stage, which focuses on learning to decode document identifiers from queries, we investigate LLM-generated queries across multiple granularity (e.g. chunks, sentences) and domain-relevant search constraints that can better capture nuanced relevancy signals. In the second stage, which aims to refine document ranking through preference learning, we explore the strategies for mining hard negatives based on the initial model's predictions. Experiments on public datasets over diverse domains demonstrate the effectiveness of our synthetic data generation and hard negative sampling approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccbffb1b-6441-4245-ab07-4c38080586f8Cited by top-tier papers2
- ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative RetrievalWeiwei Sun, Keyi Kong, Xinyu Ma, Shuaiqiang Wang et al.ICLR 2026 · 6 citations
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen et al.ICLR 2026 · 3 citations
Builds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho et al.NeurIPS 2024 · 287 citations
Related papers
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for RetrievalZhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li et al.KDD 2026 · 1 citation
- Expand, Highlight, Generate: RL-driven Document Generation for Passage RerankingArian Askari, Mohammad Aliannejadi, Chuan Meng, Evangelos Kanoulas et al.EMNLP 2023 · 7 citations
- DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG RerankersNavve Wasserman, Oliver Heinimann, Yuval Golbari, Tal Zimbalist et al.EMNLP 2025 · 7 citations
- GLEN: Generative Retrieval via Lexical Index LearningSunkyung Lee, Minjin Choi, Jongwuk LeeEMNLP 2023 · 6 citations
- DOGR: Leveraging Document-Oriented Contrastive Learning in Generative RetrievalPenghao Lu, Xin Dong, Yuansheng Zhou, Lei Cheng et al.AAAI 2025
