Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
Wenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang, Pengyue Jia, Xiaopeng Li, Yingyi Zhang, Derong Xu, Zhaocheng Du, Huifeng Guo, Ruiming Tang, Xiangyu Zhao
摘要
Retrieval-augmented generation (RAG) enhances the text generation capabilities of large language models (LLMs) by integrating external knowledge and up-to-date information. However, traditional RAG systems are limited by static workflows and lack the adaptability required for multistep reasoning and complex task management. To address these limitations, agentic RAG systems (e.g., DeepResearch) have been proposed, enabling dynamic retrieval strategies, iterative context refinement, and adaptive workflows for handling complex search queries beyond the capabilities of conventional RAG. Recent advances, such as Search-R1, have demonstrated promising gains using outcome-based reinforcement learning, where the correctness of the final answer serves as the reward signal. Nevertheless, such outcome-supervised agentic RAG methods face challenges including low exploration efficiency, gradient conflict, and sparse reward signals. To overcome these challenges, we propose to utilize fine-grained, process-level rewards to improve training stability, reduce computational costs, and enhance efficiency. Specifically, we introduce a novel method ReasonRAG that automatically constructs RAG-ProGuide, a high-quality dataset providing process-level rewards for (i) query generation, (ii) evidence extraction, and (iii) answer generation, thereby enhancing model inherent capabilities via process-supervised reinforcement learning. With the process-level policy optimization, the proposed framework empowers LLMs to autonomously invoke search, generate queries, extract relevant evidence, and produce final answers. Compared to existing approaches such as Search-R1 and traditional RAG systems, ReasonRAG, leveraging RAG-ProGuide, achieves superior performance on five benchmark datasets using only 5k training instances, significantly fewer than the 90k training instances required by Search-R1. Our code is available at https://github.com/Applied-Machine-Learning-Lab/ReasonRAG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan 等ICLR 2026 · 被引用 37 次
- SmartSearch: Process Reward-Guided Query Refinement for Search AgentsTongyu Wen, Guanting Dong, Zhicheng DouSIGIR 2026 · 被引用 13 次
- Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner CollaborationBowei He, Minda Hu, Zenan Xu, Hongru WANG 等ICML 2026 · 被引用 9 次
- Personalize Before Retrieve: LLM-based Personalized Query Expansion for User-Centric RetrievalYingyi Zhang, Pengyue Jia, Derong Xu, Yi Wen 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement LearningChuanyue Yu, Kuo Zhao, Yuhan Li, Heng Chang 等WWW 2026 · 被引用 8 次
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented GenerationPeilin Wu, Mian Zhang, Kun Wan, Wentian Zhao 等ICLR 2026 · 被引用 13 次
- CP-Search: A Chain Progressive Search Training Framework Incentivizing the Cognitive Behaviors for Searching in LLMsZehua Wang, Shipeng Li, Buzhou TangAAAI 2026
- DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented GenerationJiashuo Sun, Xianrui Zhong, Sizhe Zhou, Jiawei HanNeurIPS 2025 · 被引用 19 次
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement LearningHaoran Luo, Haihong E, Guanting Chen, Qika Lin 等ICML 2026 · 被引用 50 次
