Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
Ziyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang, Yuekai Huang, Zhiyuan Chang, Qing Wang
摘要
Information theft attacks pose a significant risk to Large Language Model (LLM) tool-learning systems. Adversaries can inject malicious commands through compromised tools, manipulating LLMs to send sensitive information to these tools, which leads to potential privacy breaches. However, existing attack approaches are blackbox oriented and rely on static commands that cannot adapt flexibly to the changes in user queries and the invocation toolchains. It makes malicious commands more likely to be detected by LLM and leads to attack failure. In this paper, we propose AUTOCMD, a dynamic attack command generation approach for information theft attacks in LLM tool-learning systems. Inspired by the concept of mimicking the familiar, AUTOCMD is capable of inferring the information utilized by upstream tools in the toolchain through learning on open-source systems and reinforcement with examples from the target systems, thereby generating more targeted commands for information theft. The evaluation results show that AUTOCMD outperforms the baselines with +13.2% ASR T hef t , and can be generalized to new tool-learning systems to expose their information leakage risks. We also design four defense methods to effectively protect tool-learning systems from the attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 被引用 37 次
- VIGIL: Defending LLM Agents Against Tool-Stream Injection via Verify-Before-CommitJunda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu 等ACL 2026 · 被引用 7 次
它引用的顶会 Paper7
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
- A First Look at Security and Privacy Risks in the RapidAPI EcosystemSong Liao, Long Cheng, Xiapu Luo, Zheng Song 等CCS 2024 · 被引用 3 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language ModelsWei Zou, Runpeng Geng, Binghui Wang, Jinyuan JiaUSENIX Security 2025
相关 Paper
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina 等CCS 2024 · 被引用 28 次
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP EcosystemShuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang 等S&P 2026 · 被引用 20 次
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 被引用 1 次
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin 等S&P 2026 · 被引用 1 次
- Endless Jailbreaks with Bijection LearningBrian R. Y. Huang, Maximilian Li, Leonard TangICLR 2025
