Towards Tool Use Alignment of Large Language Models
Zhiyuan Chen, Shiqi Shen, Guangyao Shen, Gong Zhi, Xu Chen, Yankai Lin
Abstract
Recently, tool use with LLMs has become one of the primary research topics as it can help LLM generate truthful and helpful responses.Existing studies on tool use with LLMs primarily focus on enhancing the tool-calling ability of LLMs.In practice, like chat assistants, LLMs are also required to align with human values in the context of tool use.Specifically, LLMs should refuse to answer unsafe tool use relevant instructions and insecure tool responses to ensure their reliability and harmlessness.At the same time, LLMs should demonstrate autonomy in tool use to reduce the costs associated with tool calling.To tackle this issue, we first introduce the principle that LLMs should follow in tool use scenarios: H2A.The goal of H2A is to align LLMs with helpfulness, harmlessness, and autonomy.In addition, we propose ToolAlign, a dataset comprising instruction-tuning data and preference data to align LLMs with the H2A principle for tool use.Based on ToolAlign, we develop LLMs by supervised fine-tuning and preference learning, and experimental results demonstrate that the LLMs exhibit remarkable toolcalling capabilities, while also refusing to engage with harmful content, and displaying a high degree of autonomy in tool utilization.The code and datasets are available at: https: //github.com/zhiyuanc2001/ToolAlign.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- ToolSafety: A Comprehensive Dataset for Enhancing Safety in LLM-Based Agent Tool InvocationsYuejin Xie, Youliang Yuan, Wenxuan Wang, Fan Mo et al.EMNLP 2025 · 7 citations
- AIR: Improving Agent Safety through Incident ResponseZibo Xiao, Jun Sun, Junjie ChenICML 2026 · 5 citations
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAradhye Agarwal, Gurdit Singh Siyan, Yash Pandya, Joykirat Singh et al.ICML 2026 · 5 citations
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection LearningZhiyuan Ma, Jiayu Liu, Xianzhen Luo, Zhenya Huang et al.KDD 2025 · 4 citations
- ParaTool: Shifting Tool Representations from Context to ParametersZekai Yu, Qi Meng, Qizhi Chu, Yu Hao et al.ICML 2026
Builds on14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 3 citations
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.ACL 2025
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
- Reducing Tool Hallucination via Reliability AlignmentHongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang et al.ICML 2025
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
