Alignment for Efficient Tool Calling of Large Language Models
Hongshen Xu, Zihan Wang, Zichen Zhu, Lei Pan, Xingyu Chen, Shuai Fan, Lu Chen, Kai Yu
摘要
Recent advancements in tool learning have enabled large language models (LLMs) to integrate external tools, enhancing their task performance by expanding their knowledge boundaries. However, relying on tools often introduces trade-offs between performance, speed, and cost, with LLMs sometimes exhibiting overreliance and overconfidence in tool usage. This paper addresses the challenge of aligning LLMs with their knowledge boundaries to make more intelligent decisions about tool invocation. We propose a multi-objective alignment framework that combines probabilistic knowledge boundary estimation with dynamic decision-making, allowing LLMs to better assess when to invoke tools based on their confidence. Our framework includes two methods for knowledge boundary estimation-consistency-based and absolute estimation-and two training strategies for integrating these estimates into the model's decision-making process. Experimental results on various tool invocation scenarios demonstrate the effectiveness of our framework, showing significant improvements in tool efficiency by reducing unnecessary tool usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DiSRouter: Distributed Self-Routing for LLM SelectionsHang Zheng, Hongshen Xu, Yongkai.lin, Shuai Fan 等ICLR 2026 · 被引用 6 次
- Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search TrajectoriesPeiyang Liu, Zhirui Chen, Xi Wang, Di Liang 等ACL 2026 · 被引用 4 次
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou 等KDD 2026 · 被引用 1 次
- Reducing Tool Hallucination via Reliability AlignmentHongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang 等ICML 2025
- DORA: A Dual-Objective Reinforcement Learning Framework for Effective and Efficient Multimodal Agentic SearchGuangming Qin, Yuhao Deng, Yukun Zhao, Zhenyang Li 等ACL 2026
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- Adaptive Tool Use in Large Language Models with Meta-Cognition TriggerWenjun Li, Dexun Li, Kuicai Dong, Cong Zhang 等ACL 2025
- Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary PerceptionShiyu Ni, Keping Bi, Jiafeng Guo, Lulu Yu 等ACL 2025 · 被引用 22 次
- Tool Learning in the Wild: Empowering Language Models as Automatic Tool AgentsZhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng 等WWW 2025 · 被引用 59 次
- Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy OptimizationJunyi Li, Xiaowei Qian, Yingyi Zhang, Wenlin Zhang 等ICML 2026
- MASH: Modeling Abstention via Selective Help-SeekingMustafa Omer Gul, Claire Cardie, Tanya GoyalICML 2026 · 被引用 2 次
