Tool Unlearning for Tool-Augmented LLMs
Jiali Cheng, Hadi Amiri
Abstract
Tool-augmented large language models (LLMs) are often trained on datasets of query-response pairs, which embed the ability to use tools or APIs directly into the parametric knowledge of LLMs. As these models are increasingly deployed in real-world applications, there is a need for them to forget specific tools-for example, due to security vulnerabilities, privacy regulations, or tool deprecation. This work presents "tool unlearning" as a novel machine unlearning task that presents distinct challenges beyond traditional sample-level unlearning: it requires removing functional knowledge rather than individual data points, managing the high cost of LLM optimization, and developing principled evaluation metrics. To address these challenges, we propose TOOLDELETE, the first unlearning framework designed specifically for tool-augmented LLMs. It implements three key properties for effective tool unlearning and introduces a new membership inference attack (MIA) model for effective evaluation. Extensive experiments on multiple tool learning datasets and tool-augmented LLMs show that TOOLDELETE effectively unlearns both randomly selected and class-specific tools, while preserving knowledge on remaining tools and maintaining performance on general tasks. (a) Tool Learning and Tool Unlearning Tool Deletion Requests (Insecure tools, Broken tools, ...) (c) ToolDelete (b) Traditional Unlearning vs. Tool Unlearning
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
Related papers
- A Reliable Cryptographic Framework for Empirical Machine Unlearning EvaluationYiwen Tu, Pingbang Hu, Jiaqi MaNeurIPS 2025 · 6 citations
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Decoding-Unlearning: Fact Forgetting via Entropy-Guided InferenceJingwen Pu, Mingjun Shi, Xinrui Ren, Yizhe Wang et al.ACL 2026
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang et al.AAAI 2025 · 17 citations
- Large Language Model Unlearning for Source CodeXue Jiang, Yihong Dong, Huangzhao Zhang, Tangxinyu Wang et al.AAAI 2026
