Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning System
Ziyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang, Yuekai Huang, Zhiyuan Chang, Qing Wang
Abstract
Information theft attacks pose a significant risk to Large Language Model (LLM) tool-learning systems. Adversaries can inject malicious commands through compromised tools, manipulating LLMs to send sensitive information to these tools, which leads to potential privacy breaches. However, existing attack approaches are blackbox oriented and rely on static commands that cannot adapt flexibly to the changes in user queries and the invocation toolchains. It makes malicious commands more likely to be detected by LLM and leads to attack failure. In this paper, we propose AUTOCMD, a dynamic attack command generation approach for information theft attacks in LLM tool-learning systems. Inspired by the concept of mimicking the familiar, AUTOCMD is capable of inferring the information utilized by upstream tools in the toolchain through learning on open-source systems and reinforcement with examples from the target systems, thereby generating more targeted commands for information theft. The evaluation results show that AUTOCMD outperforms the baselines with +13.2% ASR T hef t , and can be generalized to new tool-learning systems to expose their information leakage risks. We also design four defense methods to effectively protect tool-learning systems from the attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68b3e647-d9b7-43f5-a3f4-ede8e7292b70Cited by top-tier papers2
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 37 citations
- VIGIL: Defending LLM Agents Against Tool-Stream Injection via Verify-Before-CommitJunda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu et al.ACL 2026 · 7 citations
Builds on7
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- A First Look at Security and Privacy Risks in the RapidAPI EcosystemSong Liao, Long Cheng, Xiapu Luo, Zheng Song et al.CCS 2024 · 3 citations
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language ModelsWei Zou, Runpeng Geng, Binghui Wang, Jinyuan JiaUSENIX Security 2025
Related papers
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP EcosystemShuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang et al.S&P 2026 · 20 citations
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin et al.S&P 2026 · 1 citation
- Endless Jailbreaks with Bijection LearningBrian R. Y. Huang, Maximilian Li, Leonard TangICLR 2025
