Agentic Reinforcement Learning with Implicit Step Rewards
Xiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang, Yongbin Li, Jianbin Jiao, Junge Zhang
Abstract
Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, sparse and sometimes unverifiable rewards make it extremely challenging to assign credit when training LLM agents that serve as a policy. Recent work attempts to integrate process supervision into RL but suffers from biased annotation, reward hacking, high-variance from overly fine-grained rewards or failures when state overlap is rare. We therefore introduce implicit step rewards for agentic RL (iStar), a general credit-assignment strategy that integrates seamlessly with standard RL algorithms without relying on additional rollouts or explicit step labels. Particularly, we alternatively optimize an implicit process reward model (PRM) with the policy model to generate step rewards for each action via a multi-turn DPO objective. Theoretical analysis shows that this learning objective produces a step-wise reward function learned from trajectory preferences. Then the implicit step rewards are used to compute step-level advantages, which are combined with trajectory (or episode)-level advantages for policy updates, creating a self-reinforcing training loop. We evaluate our method on three challenging agent benchmarks, including WebShop and VisualSokoban, as well as open-ended social interactions with unverifiable rewards in SOTOPIA. Crucially, our method shows superior performance over frontier LLMs and strong RL baselines across domains, achieving state-of-the-art results with higher sample-efficiency and training stability. Further analysis also demonstrates efficient exploration by iStar with increased rewards in both step- and episode-level while maintaining fewer steps to achieve task success.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac11461d-d00f-466e-8b60-d249ce782c1bCited by top-tier papers4
- Adaptive Social Learning via Mode Policy Optimization for Language AgentsMinzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang et al.ICLR 2026 · 15 citations
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li et al.ICLR 2026 · 9 citations
- Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsShuai Zhen, Yanhua Yu, Ruopei Guo, Nan Cheng et al.ACL 2026 · 2 citations
- HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM AgentsJiangweizhi Peng, Yuanxin Liu, Ruida Zhou, Charles Fleming et al.ICML 2026
Builds on27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Self-evolving LLM agents with in-distribution OptimizationYudi Zhang, Meng Fang, Zhenfang Chen, Mykola PechenizkiyICML 2026 · 1 citation
- Watch Every Step! LLM Agent Learning via Iterative Step-level Process RefinementWeimin Xiong, Yifan Song, Xiutian Zhao, Wenhao Wu et al.EMNLP 2024 · 8 citations
- Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement LearningCheng Xin, Shuo He, Lang Feng, Haiyang Xu et al.ICML 2026 · 6 citations
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan et al.ICLR 2026 · 37 citations
