Reinforcing Agentic Search Via Reward Density Optimization
Kun Luo, Hongjin Qian, Zheng Liu, Ziyi Xia, Shitao Xiao, Zhao Cao, Siqi Bao, Jun Zhao, Kang Liu
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is a promising approach for enhancing agentic search. However, its performance is often hindered by reward sparsity, whereby agents receive very limited positive feedback despite incurring significant exploration costs. In this paper, we formalize this challenge as a new research problem termed Reward Density Optimization, which aims to improve the reward obtained per unit of exploration cost. To address this problem, we introduce InfoFlow, a systematic framework that operates along three complementary dimensions: 1) Sub-goal Scaffolding: which decomposes long-horizon tasks into intermediate objectives and assigns process-level rewards to provide denser learning signals; 2) Pathfinding Hints: which injects corrective guidance into stalled trajectories to increase the ratio of successful trials; and 3) Dual-agent Refinement: which employs a dual-agent architecture to offload the cognitive burden of deep exploration. We evaluate InfoFlow on several popular agentic search benchmarks, where it significantly outperforms strong baselines and enables lightweight LLMs to achieve performance comparable to that of advanced proprietary models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Diversity-Incentivized Exploration for Versatile ReasoningZican Hu, Shilin Zhang, Yafu Li, Jianhao Yan et al.ICLR 2026 · 32 citations
- In-the-Flow Agentic System Optimization for Effective Planning and Tool UseZhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu et al.ICLR 2026 · 65 citations
- DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Tree-based SearchFang Wu, Weihao Xuan, Heli Qi, Aaron Tu et al.ICLR 2026 · 6 citations
- See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language ModelsYin Zhang, Zonghan Wu, Jiaxuan Zhao, Junfeng Fang et al.ICML 2026
- Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language ModelsYan Liu, Feng Zhang, Zhanyu Ma, Jun Xu et al.ACL 2026 · 2 citations
