On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
Tong Liu, Cheng Qian, Matej Cief, Yuan He, Daniele Dan, Nikolaos Aletras, Gabriella Kazai
Abstract
Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along two complementary axes: effectiveness , i.e., how this capability is measured , and efficiency , i.e., how it is learned . On effectiveness, we systematically analyze tool-calling evaluation pipelines and show that results can be highly sensitive to seemingly minor, often undocumented implementation choices including the random seed , system prompt , multi-turn template construction , and how prior interaction/reasoning history is carried forward. These choices can lead to substantial differences in reported performance, especially in multi-turn settings where without rigorous standardization, leaderboard rankings are unreliable. On efficiency, we examine standard reinforcement learning (RL) for tool-calling and identify two sources of computational waste: (i) during rollouts, many prompts produce no learning signal, and (ii) during policy updates, optimization incurs high computational cost. Guided by these findings, we introduce two techniques that accelerate RL-based tool-calling training, achieving substantial wall-clock speedup without degrading performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89deb826-2ecf-4cce-beff-0f8185f790caBuilds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- Empowering LLM Tool Invocation with Tool-call Reward ModelDa Ma, Ziyue Yang, Hongshen Xu, Haotian Fang et al.ICLR 2026
- ToolBox-RL: Learning to Generalize Tool Use Across Massive RepositoriesXinyan Shi, Renzhi Wang, Haodong Liu, Piji LiWWW 2026
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured TemplatesHy Dang, Tianyi Liu, Zhuofeng Wu, Jingfeng Yang et al.EMNLP 2025
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced ReasoningShaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz et al.ICLR 2026 · 61 citations
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function CallingJianghao Lin, Yuanyuan Shi, Xin Peng, Renjie Ding et al.ACL 2026 · 3 citations
