AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
Zhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou, Jian Guan, Wei Wu, Tao Ji, Tao Gui
Abstract
Despite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50b1bf49-4647-401b-be5f-ce71b6faa705Cited by top-tier papers4
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL SystemYinjie Wang, Tianbao Xie, Ke Shen, Mengdi Wang et al.ICML 2026 · 15 citations
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive RewardsShiyu Li, Yifan Wang, Peiming Li, Zheng Wei et al.ICML 2026 · 8 citations
- GLARE: Scalable Neuro-Symbolic Reward Shaping for LLM Agents via Group-Level AutomataJingyuan Yan, Qingchen Liu, Qichao Ma, Jiahu QinICML 2026
- AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RLZhiheng Xi, Jixuan Huang, Chenyang Liao, Baodai Huang et al.ICLR 2026
Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
Related papers
- Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningAmrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng et al.ICLR 2025
- Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsHyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim et al.NeurIPS 2025 · 37 citations
- DEPO: Dual-Efficiency Preference Optimization for LLM AgentsSirui Chen, Mengshi Zhao, Lei Xu, Yuying Zhao et al.AAAI 2026 · 2 citations
- Scaling Autonomous Agents via Automatic Reward Modeling And PlanningZhenfang Chen, Delin Chen, Rui Sun, Wenjun Liu et al.ICLR 2025
- Dynamic and Generalizable Process Reward ModelingZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng et al.ACL 2025 · 13 citations
