Lune

ICML2026Top-tier venue

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards

Shiyu Li, Yifan Wang, Peiming Li, Zheng Wei, Yang Tang

2026Year
8Citations

Abstract

Search agents powered by Large Language Models (LLMs) have demonstrated significant potential in tackling knowledge-intensive tasks. Reinforcement learning (RL) has emerged as a powerful paradigm for training these agents to perform complex, multi-step reasoning. However, prior RL-based methods often rely on sparse or rule-based rewards, which can lead agents to commit to suboptimal or erroneous reasoning paths without the ability to recover. To address these limitations, we propose ReSeek, a self-correcting framework enabling search agents to recover from erroneous search paths during an episode. By invoking a special JUDGE action, the agent can judge the information and re-plan its search strategy. To guide this process, we design a dense, instructive process reward function, which combines an answer correctness reward for task completion with a self-correction reward that trains the agent to judge the utility of retrieved information. Additionally, to mitigate the risk of data contamination in existing datasets, we introduce FictionalHot, a contamination-resistant benchmark requiring complex reasoning. Experiments show ReSeek significantly outperforms SOTA baselines in task success and path faithfulness. Our code and dataset are available at https: //github.com/TencentBAC/ReSeek .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 28dd8553-e00d-48b1-a57f-dd93cb35296a

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines