Agentic Reinforced Policy Optimization
Guanting Dong, Hangyu Mao, Kai Ma, Licheng Bao, Yifei Chen, Zhongyuan Wang, Zhongxia Chen, Jiazhen Du, Huiyang Wang, Fuzheng Zhang, Guorui Zhou, Yutao Zhu
Abstract
Large-scale reinforcement learning with verifiable rewards (RLVR) has demonstrated its effectiveness in harnessing the potential of large language models (LLMs) for single-turn reasoning tasks. In realistic reasoning scenarios, LLMs can often utilize external tools to assist in task-solving processes. However, current RL algorithms inadequately balance the models'intrinsic long-horizon reasoning capabilities and their proficiency in multi-turn tool interactions. To bridge this gap, we propose Agentic Reinforced Policy Optimization (ARPO), a novel agentic RL algorithm tailored for training multi-turn LLM-based agents. Through preliminary experiments, we observe that LLMs tend to exhibit highly uncertain behavior, characterized by an increase in the entropy distribution of generated tokens, immediately following interactions with external tools. Motivated by this observation, ARPO incorporates an entropy-based adaptive rollout mechanism, dynamically balancing global trajectory sampling and step-level sampling, thereby promoting exploration at steps with high uncertainty after tool usage. By integrating an advantage attribution estimation, ARPO enables LLMs to internalize advantage differences in stepwise tool-use interactions. Our experiments across 13 challenging benchmarks in computational reasoning, knowledge reasoning, and deep search domains demonstrate ARPO's superiority over trajectory-level RL algorithms. Remarkably, ARPO achieves improved performance using only half of the tool-use budget required by existing methods, offering a scalable solution for aligning LLM-based agents with real-time dynamic environments. Our code and datasets are released at https://github.com/dongguanting/ARPO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers51
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningZhenghai Xue, Longtao Zheng, Qian Liu, Yingru Li et al.ICLR 2026 · 152 citations
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen et al.ICLR 2026 · 71 citations
- TreePO: Enhancing Policy Efficacy and Inference Efficiency with Tree ModelingYizhi Li, Qingshui Gu, Zhoufutu Wen, Ziniu Li et al.ICML 2026 · 61 citations
- OneThinker: All-in-one Reasoning Model for Image and VideoKaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan et al.CVPR 2026 · 55 citations
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI AgentsYifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu et al.ICLR 2026 · 45 citations
Builds on32
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
Related papers
- AT²PO: Agentic Turn-based Policy Optimization via Tree SearchZefang Zong, Dingwei Chen, Yang Li, Qi Yi et al.ACL 2026 · 3 citations
- Toward Generalized Web Agent Training: A Deep Dive into Entropy-Balanced Reinforcement LearningGuanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao et al.WWW 2026 · 2 citations
- RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy OptimizationSiwei Zhang, Yun Xiong, Xi Chen, Zian Jia et al.KDD 2026 · 7 citations
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan et al.ICLR 2026 · 37 citations
- Empowering LLM Tool Invocation with Tool-call Reward ModelDa Ma, Ziyue Yang, Hongshen Xu, Haotian Fang et al.ICLR 2026
