Lune

ACL2025顶会

Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System

Ziyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang, Yuekai Huang, Zhiyuan Chang, Qing Wang

2025年份
2顶会引用

摘要

Information theft attacks pose a significant risk to Large Language Model (LLM) tool-learning systems. Adversaries can inject malicious commands through compromised tools, manipulating LLMs to send sensitive information to these tools, which leads to potential privacy breaches. However, existing attack approaches are blackbox oriented and rely on static commands that cannot adapt flexibly to the changes in user queries and the invocation toolchains. It makes malicious commands more likely to be detected by LLM and leads to attack failure. In this paper, we propose AUTOCMD, a dynamic attack command generation approach for information theft attacks in LLM tool-learning systems. Inspired by the concept of mimicking the familiar, AUTOCMD is capable of inferring the information utilized by upstream tools in the toolchain through learning on open-source systems and reinforcement with examples from the target systems, thereby generating more targeted commands for information theft. The evaluation results show that AUTOCMD outperforms the baselines with +13.2% ASR T hef t , and can be generalized to new tool-learning systems to expose their information leakage risks. We also design four defense methods to effectively protect tool-learning systems from the attack.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖