LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
Kangning Zhang, Weiwen Liu, Wenxiang Jiao, Kounianhua Du, Yuan Lu, Weinan Zhang, Yong Yu
Abstract
Augmenting Large Language Models (LLMs) with external tools enables them to execute complex, multi-step tasks. However, tool learning is hampered by the static synthetic data pipelines where data generation and model training are executed as two separate, non-interactive processes. This approach fails to adaptively focus on a model's specific weaknesses and allows noisy labels to persist, degrading training efficiency. We introduce LoopTool, a fully automated, model-aware data evolution framework that closes this loop by tightly integrating data synthesis and model training. LoopTool iteratively refines both the data and the model through three synergistic modules: (1) Greedy Capability Probing (GCP) diagnoses the model's mastered and failed capabilities; (2) Judgement-Guided Label Verification (JGLV) uses an open-source judge model to find and correct annotation errors, progressively purifying the dataset; and (3) Error-Driven Data Expansion (EDDE) generates new, challenging samples based on identified failures. This closed-loop process operates within a cost-effective, open-source ecosystem, eliminating dependence on expensive closed-source APIs. Experiments show that our 8B model trained with LoopTool significantly surpasses its 32B data generator and achieves new state-of-the-art results on the BFCL-v3 and ACEBench benchmarks for its scale. Our work demonstrates that closed-loop, self-refining data pipelines can dramatically enhance the tool-use capabilities of LLMs. 1 * This work was done while Kangning Zhang and Kounianhua Du were interns at Xiaohongshu Inc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 867d1fa7-d61e-411b-8a49-61c06bb4db52Cited by top-tier papers5
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
- Robust Tool Use via Fission-GRPO: Learning to Recover from Execution ErrorsZhiwei Zhang, Fei Zhao, Rui Wang, Zezhong Wang et al.ACL 2026 · 4 citations
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun et al.CHI 2026 · 2 citations
- A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and SolutionsZhiyin Yu, Yuchen Mou, Juncheng Yan, Junyu Luo et al.ACL 2026
- A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and UsageCongmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen et al.ACL 2026
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
Related papers
- ToolACE: Winning the Points of LLM Function CallingWeiwen Liu, Xu Huang, Xingshan Zeng, Xinlong Hao et al.ICLR 2025
- ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learningXingshan Zeng, Weiwen Liu, Xu Huang, Zezhong Wang et al.AAAI 2026 · 3 citations
- iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool UseYirong Zeng, Xiao Ding, Yuxian Wang, Weiwen Liu et al.EMNLP 2025
- From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal ModelsHongrui Jia, Chaoya Jiang, Yongrui Heng, Shikun Zhang et al.ICML 2026
- SIPDO: Closed-Loop Prompt Optimization via Synthetic Data FeedbackYaoning Yu, Ye Yu, Peiyan Zhang, Kai Wei et al.ICLR 2026 · 7 citations
