Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
Yuchong Xie, Mingyu Luo, Zesen Liu, Zhixiang Zhang, Kaikai Zhang, Yu Liu, Ci Tao, Changhui Wang, Zongjie Li, Ping Chen, Shuai Wang, Dongdong She
Abstract
Coding agents powered by large language models are becoming central modules of modern IDEs. They help users to perform various complex coding tasks by invoking tools. Although powerful, tool-invocation operation in coding agents opens a substantial attack surface for adversaries. Prior work has demonstrated attacks against both general-purpose LLM agents and domain-specific agents. However, to our knowledge, no previous works focus on the security risks of tool-invocation in coding agents. To fill this gap, we conduct the first systematic, in-depth red-teaming from a tool-invocation perspective in six popular real-world coding agents: Cursor, Claude Code, Copilot, Windsurf, Cline, and Trae. Our red-teaming proceeds in two phases. In Phase 1, we conduct a prompt leakage as reconnaissance to recover system prompts and related context. Specifically, we identify a mode gap between chat generation and tool-call argument generation: during schema-driven argument completion, the model behaves as if it is performing benign structured filling and may copy hidden agent context into tool-call arguments. We instantiate this gap as ToolLeak, which exfiltrates agent-internal prompts (e.g., system prompts and tool metadata) via required tool parameters. In Phase 2, we hijack the tool-invocation behavior of the coding agent with a novel two-channel prompt injection in the tool description and the tool return. Our hijacking achieves remote code execution (RCE) on major real-world coding agents. We adaptively construct the malicious payload using leaked security information in Phase 1. Our evaluation shows that ToolLeak substantially outperforms strong prompt-leak baselines in both emulated and real-world settings. In the emulated setting, ToolLeak achieves the best overall prompt-exfiltration performance across all six simulated coding agents. On real-world coding agents, ToolLeak achieves the best pseudo-recall on 18 of 25 evaluated agent-LLM pairs. Furthermore, our red-teaming successfully hijacks all six real-world coding agents for RCE and consistently yields higher attack success rates than baseline attacks. Lastly, we present two case studies on Cursor and Claude Code to demonstrate the real-world impact of our red-teaming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bd1d4d1-1773-4896-bbec-4fd2acea2899Builds on19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou et al.NDSS 2026 · 181 citations
Related papers
- SOPE: Situation-Aware and Statistically Indistinguishable Privacy Exfiltration for MCP-enabled AgentsRuixiao Lin, Qingming Li, Jiahao Chen, Chunyi Zhou et al.ICML 2026
- Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMsXiang Zheng, YUTAO WU, Hanxun Huang, Yige Li et al.ICML 2026
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 37 citations
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP EcosystemShuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang et al.S&P 2026 · 20 citations
- Black-Box Adversarial Attacks on LLM-Based Code CompletionSlobodan Jenko, Niels Mündler, Jingxuan He, Mark Vero et al.ICML 2025
