Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning
Guanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Yutao Zhu, Hangyu Mao, Guorui Zhou, Zhicheng Dou, Ji-Rong Wen
摘要
Recently, Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities through reinforcement learning (RL). However, enabling LLM-based agents to effectively orchestrate multiple tools remains an open challenge. In this paper, we introduce Tool-Star, an end-to-end agentic post-training framework that empowers LLM-based web agents to strategically interact with external multi-tool environments. Tool-Star begins with a general tool-integrated data synthesis pipeline that combines two complementary sampling strategies to generate tool-use trajectories, followed by quality normalization and difficulty-aware curriculum construction to filter noisy samples and organize training data from easy to hard. We then introduce a two-stage training paradigm for multi-tool collaborative reasoning: (1) cold-start supervised fine-tuning with tool feedback to bootstrap long-horizon tool-augmented reasoning, and (2) a multi-tool self-critic RL algorithm with hierarchical rewards to reinforce effective tool coordination. Experiments across 13 benchmarks demonstrate Tool-Star's effectiveness. Further analyses provide practical insights for optimizing strategic tool use in web agents. The code is available at https://github.com/RUC-NLPIR/Tool-Star.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical ReasoningQikai Chang, Zhenrong Zhang, Pengfei Hu, Jun Du 等ICLR 2026 · 被引用 8 次
- ToolBox-RL: Learning to Generalize Tool Use Across Massive RepositoriesXinyan Shi, Renzhi Wang, Haodong Liu, Piji LiWWW 2026
- ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded ExecutionShouzheng Huang, Meishan Zhang, Baotian Hu, Min ZhangACL 2026
- Small LLMs Are Weak Tool Learners: A Multi-LLM AgentWeizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan 等EMNLP 2024 · 被引用 18 次
