ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
Shiyu Li, Yifan Wang, Peiming Li, Zheng Wei, Yang Tang
Abstract
Search agents powered by Large Language Models (LLMs) have demonstrated significant potential in tackling knowledge-intensive tasks. Reinforcement learning (RL) has emerged as a powerful paradigm for training these agents to perform complex, multi-step reasoning. However, prior RL-based methods often rely on sparse or rule-based rewards, which can lead agents to commit to suboptimal or erroneous reasoning paths without the ability to recover. To address these limitations, we propose ReSeek, a self-correcting framework enabling search agents to recover from erroneous search paths during an episode. By invoking a special JUDGE action, the agent can judge the information and re-plan its search strategy. To guide this process, we design a dense, instructive process reward function, which combines an answer correctness reward for task completion with a self-correction reward that trains the agent to judge the utility of retrieved information. Additionally, to mitigate the risk of data contamination in existing datasets, we introduce FictionalHot, a contamination-resistant benchmark requiring complex reasoning. Experiments show ReSeek significantly outperforms SOTA baselines in task success and path faithfulness. Our code and dataset are available at https: //github.com/TencentBAC/ReSeek .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28dd8553-e00d-48b1-a57f-dd93cb35296aBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
- Iterative Self-Incentivization Empowers Large Language Models as Agentic SearchersZhengliang Shi, Lingyong Yan, Dawei Yin, Suzan Verberne et al.NeurIPS 2025 · 15 citations
- LeTS: Learning to Think-and-Search via Process-and-Outcome Reward HybridizationQi Zhang, Shouqing Yang, Lirong Gao, Hao Chen et al.EMNLP 2025
- Do LLM Agents Know How to Ground, Recover, and Assess? Evaluating Epistemic Competence in Information-Seeking AgentsJiaqi Shao, Yuxiang Lin, Munish Prasad Lohani, Yufeng Miao et al.ICLR 2026 · 4 citations
- Open Data Synthesis for Deep ResearchZiyi Xia, Kun Luo, Hongjin Qian, Siqi Bao et al.ICLR 2026 · 14 citations
