Towards Tool Use Alignment of Large Language Models
Zhiyuan Chen, Shiqi Shen, Guangyao Shen, Gong Zhi, Xu Chen, Yankai Lin
摘要
Recently, tool use with LLMs has become one of the primary research topics as it can help LLM generate truthful and helpful responses.Existing studies on tool use with LLMs primarily focus on enhancing the tool-calling ability of LLMs.In practice, like chat assistants, LLMs are also required to align with human values in the context of tool use.Specifically, LLMs should refuse to answer unsafe tool use relevant instructions and insecure tool responses to ensure their reliability and harmlessness.At the same time, LLMs should demonstrate autonomy in tool use to reduce the costs associated with tool calling.To tackle this issue, we first introduce the principle that LLMs should follow in tool use scenarios: H2A.The goal of H2A is to align LLMs with helpfulness, harmlessness, and autonomy.In addition, we propose ToolAlign, a dataset comprising instruction-tuning data and preference data to align LLMs with the H2A principle for tool use.Based on ToolAlign, we develop LLMs by supervised fine-tuning and preference learning, and experimental results demonstrate that the LLMs exhibit remarkable toolcalling capabilities, while also refusing to engage with harmful content, and displaying a high degree of autonomy in tool utilization.The code and datasets are available at: https: //github.com/zhiyuanc2001/ToolAlign.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ToolSafety: A Comprehensive Dataset for Enhancing Safety in LLM-Based Agent Tool InvocationsYuejin Xie, Youliang Yuan, Wenxuan Wang, Fan Mo 等EMNLP 2025 · 被引用 7 次
- AIR: Improving Agent Safety through Incident ResponseZibo Xiao, Jun Sun, Junjie ChenICML 2026 · 被引用 5 次
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAradhye Agarwal, Gurdit Singh Siyan, Yash Pandya, Joykirat Singh 等ICML 2026 · 被引用 5 次
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection LearningZhiyuan Ma, Jiayu Liu, Xianzhen Luo, Zhenya Huang 等KDD 2025 · 被引用 4 次
- ParaTool: Shifting Tool Representations from Context to ParametersZekai Yu, Qi Meng, Qizhi Chu, Yu Hao 等ICML 2026
它引用的顶会 Paper14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 被引用 3 次
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 等ACL 2025
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang 等NeurIPS 2023 · 被引用 463 次
- Reducing Tool Hallucination via Reliability AlignmentHongshen Xu, Zichen Zhu, Lei Pan, Zihan Wang 等ICML 2025
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen 等FSE 2025 · 被引用 1 次
