RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
Xinze Li, Sen Mei, Zhenghao Liu, Yukun Yan, Shuo Wang, Shi Yu, Zheni Zeng, Hao Chen, Ge Yu, Zhiyuan Liu, Maosong Sun, Chenyan Xiong
Abstract
Retrieval-Augmented Generation (RAG) has proven its effectiveness in mitigating hallucinations in Large Language Models (LLMs) by retrieving knowledge from external resources. To adapt LLMs for the RAG systems, current approaches use instruction tuning to optimize LLMs, improving their ability to utilize retrieved knowledge. This supervised fine-tuning (SFT) approach focuses on equipping LLMs to handle diverse RAG tasks using different instructions. However, it trains RAG modules to overfit training signals and overlooks the varying data preferences among agents within the RAG system. In this paper, we propose a Differentiable Data Rewards (DDR) method, which end-to-end trains RAG systems by aligning data preferences between different RAG modules. DDR works by collecting the rewards to optimize each agent in the RAG system with the rollout method, which prompts agents to sample some potential responses as perturbations, evaluates the impact of these perturbations on the whole RAG system, and subsequently optimizes the agent to produce outputs that improve the performance of the RAG system. Our experiments on various knowledge-intensive tasks demonstrate that DDR significantly outperforms the SFT method, particularly for LLMs with smaller-scale parameters that depend more on the retrieved knowledge. Additionally, DDR exhibits a stronger capability to align the data preference between RAG modules. The DDR method makes the generation module more effective in extracting key information from documents and mitigating conflicts between parametric memory and external knowledge. All codes are available at https://github.com/OpenMatch/RAG-DDR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51de625a-047f-4141-b501-41e9c62a3f7eCited by top-tier papers12
- Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement LearningYiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma et al.NeurIPS 2025 · 47 citations
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented ReasoningYaorui Shi, Sihang Li, Chang Wu, Zhiyuan Liu et al.NeurIPS 2025 · 30 citations
- RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-ThoughtsMingyan Wu, Zhenghao Liu, Yukun Yan, Xinze Li et al.ACL 2025 · 16 citations
- Iterative Self-Incentivization Empowers Large Language Models as Agentic SearchersZhengliang Shi, Lingyong Yan, Dawei Yin, Suzan Verberne et al.NeurIPS 2025 · 15 citations
- ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented GenerationPengcheng Huang, Zhenghao Liu, Yukun Yan, Haiyan Zhao et al.NeurIPS 2025 · 11 citations
Builds on21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
Related papers
- Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented GenerationGuanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang et al.WWW 2025 · 44 citations
- Boosting the Potential of Large Language Models with an Intelligent Information AssistantYujia Zhou, Zheng Liu, Zhicheng DouNeurIPS 2024
- ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented GeneratorJunda Zhu, Lingyong Yan, Haibo Shi, Dawei Yin et al.EMNLP 2024 · 6 citations
- KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language ModelsRuizhe Zhang, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu et al.AAAI 2025 · 15 citations
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement LearningHaoran Luo, Haihong E, Guanting Chen, Qika Lin et al.ICML 2026 · 50 citations
