Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
Kanghua Mo, Li Hu, Yucheng Long, Zhihao Li
摘要
Large language model (LLM) agents have demonstrated remarkable capabilities in complex reasoning and decision-making by leveraging external tools. However, this tool-centric paradigm introduces a previously underexplored attack surface, where adversaries can manipulate tool metadata -- such as names, descriptions, and parameter schemas -- to influence agent behavior. We identify this as a new and stealthy threat surface that allows malicious tools to be preferentially selected by LLM agents, without requiring prompt injection or access to model internals. To demonstrate and exploit this vulnerability, we propose the Attractive Metadata Attack (AMA), a black-box in-context learning framework that generates highly attractive but syntactically and semantically valid tool metadata through iterative optimization. The proposed attack integrates seamlessly into standard tool ecosystems and requires no modification to the agent's execution framework. Extensive experiments across ten realistic, simulated tool-use scenarios and a range of popular LLM agents demonstrate consistently high attack success rates (81%-95%) and significant privacy leakage, with negligible impact on primary task execution. Moreover, the attack remains effective even against prompt-level defenses, auditor-based detection, and structured tool-selection protocols such as the Model Context Protocol, revealing systemic vulnerabilities in current agent architectures. These findings reveal that metadata manipulation constitutes a potent and stealthy attack surface. Notably, AMA is orthogonal to injection attacks and can be combined with them to achieve stronger attack efficacy, highlighting the need for execution-level defenses beyond prompt-level and auditor-based mechanisms. Code is available at https://github.com/SEAIC-M/AMA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language ModelsThierry Blankenstein, Jialin Yu, Zixuan Li, Vassilis Plachouras 等ICLR 2026 · 被引用 8 次
- Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated PlanningShanghao Shi, Xiao Wang, Chaoyu Zhang, Hao Li 等ICML 2026 · 被引用 1 次
- Conjunctive Prompt Attacks in Multi-Agent LLM SystemsNokimul Hasan Arif, Qian Lou, Mengxin ZhengACL 2026
它引用的顶会 Paper7
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksVaidehi Patil, Peter Hase, Mohit BansalICLR 2024 · 被引用 167 次
- Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction AmplificationBoyang Zhang, Yicong Tan, Yun Shen, Ahmed Salem 等EMNLP 2025 · 被引用 4 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
- Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning SystemZiyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang 等ACL 2025
相关 Paper
- MPMA: Preference Manipulation Attack Against Model Context ProtocolZihan Wang, Rui Zhang, Yu Liu, Wenshu Fan 等AAAI 2026 · 被引用 29 次
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang 等ACL 2025
- Tool Preferences in Agentic LLMs are UnreliableKazem Faghih, Wenxiao Wang, Yize Cheng, Siddhant Bharti 等EMNLP 2025
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou 等NDSS 2026 · 被引用 181 次
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP EcosystemShuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang 等S&P 2026 · 被引用 20 次
