ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai, Guojun Yin, Wei Lin, Ran He
Abstract
Large language models (LLMs) transcend passive generation and act as goal-directed agents by invoking external tools. Reinforcement learning (RL) offers a principled framework for optimizing these emergent tool-use policies, yet the prevailing paradigm relies exclusively on sparse outcome rewards and lacks consideration of the particularity of tool-use tasks, inflating policy-gradient variance and resulting in inefficient training. To better understand and address these challenges, we first establish a theoretical link between policy entropy and training stability of tool-use tasks, which reveals that structured, low-entropy tokens are primary determinants of rewards. Motivated by this insight, we propose Reshaped Token-level policy gradients (ResT) for tool-use tasks. ResT reshapes the policy gradient through entropy-informed token reweighting, progressively upweighting reasoning tokens as training proceeds. This entropy-aware scheme enables a smooth shift from structural correctness to semantic reasoning and stabilizes convergence in multi-turn tool-use tasks. Evaluation on BFCL and API-Bank shows that ResT achieves state-of-the-art results, outperforming prior methods by up to . When fine-tuned on a 4B base LLM, ResT further surpasses GPT-4o by on single-turn tasks and on multi-turn base tasks. Code is available at https://github.com/1229095296/ResT_Tool_use_LLM.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae0b598b-40f9-4515-b861-95bf3f152d2eCited by top-tier papers1
Ask how each one uses itBuilds on13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
Related papers
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
- Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model AgentsZeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen et al.ACL 2026
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao et al.ICLR 2026 · 146 citations
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
