RPDR: A Round-trip Prediction-Based Data Augmentation Framework for Long-Tail Question Answering
Yiming Zhang, Siyue Zhang, Junbo Zhao, Chen Zhao
Abstract
Long-tail question answering presents significant challenges for large language models (LLMs) due to their limited ability to acquire and accurately recall less common knowledge. Retrieval-augmented generation (RAG) systems have shown great promise in mitigating this limitation by integrating external retrieval mechanisms. However, dense retrieval models often face the same difficulties when generalizing to rare or niche knowledge. In this study, we introduce RPDR, a novel data augmentation framework that selects high-quality easy-to-learn training data, to enhance dense retrievers. Our approach is built around three core components: synthetic data generation, data selection with Round-Trip prediction to identify easy-to-learn instances, and retriever training with these instances. We evaluate RPDR on two longtail retrieval benchmarks, POPQA and ENTI-TYQUESTIONS, demonstrating substantial improvements over existing retrievers like BM25 and Contriver, especially on extremely longtail categories. We identify the strengths and limitations of RPDR through detailed human analysis and propose a dynamic routing mechanism to dynamically route queries to specialized retrieval modules to further improve retrieval performance. 1 * Work done when Yiming Zhang was visiting NYU Shanghai.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f367dd4-46f3-487b-9d4b-1ec27e15b8a5Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsYue Yu, Wei Ping, Zihan Liu, Boxin Wang et al.NeurIPS 2024 · 321 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
Related papers
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval AugmentationHongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao et al.WWW 2025 · 92 citations
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question AnsweringLei Li, Xiao Zhou, Yingying Zhang, Xian WuWWW 2026
- LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question AnsweringQingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha et al.EMNLP 2024 · 13 citations
- Boosting Retrieval-Augmented Generation with Generation-Augmented Retrieval: A Co-Training ApproachYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented InstructionsWanlong Liu, Junying Chen, Ke Ji, Li Zhou et al.EMNLP 2025 · 1 citation
