Agents Under Siege: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt Attacks
Rana Muhammad Shahroz, Zhen Tan, Sukwon Yun, Charles Fleming, Tianlong Chen
Abstract
Most discussions about Large Language Model (LLM) safety have focused on single-agent settings but multi-agent LLM systems now create novel adversarial risks because their behavior depends on communication between agents and decentralized reasoning. In this work, we innovatively focus on attacking pragmatic systems that have constrains such as limited token bandwidth, latency between message delivery, and defense mechanisms. We design a that optimizes prompt distribution across latency and bandwidth-constraint network topologies to bypass distributed safety mechanisms within the system. Formulating the attack path as a problem of , coupled with the novel , we leverage graph-based optimization to maximize attack success rate while minimizing detection risk. Evaluating across models including , , , and other variants on various datasets like and , our method outperforms conventional attacks by up to , exposing critical vulnerabilities in multi-agent systems. Moreover, we demonstrate that existing defenses, including variants of and , fail to prohibit our attack, emphasizing the urgent need for multi-agent specific safety mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- SAGA: A Security Architecture for Governing AI Agentic SystemsGeorgios Syros, Anshuman Suri, Jacob Ginesin, Cristina Nita-Rotaru et al.NDSS 2026 · 63 citations
- When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent SystemsLingxi Zhang, Guangtao Zheng, Hanjie ChenICML 2026 · 1 citation
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang et al.ICML 2026
- A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent SystemsTianhao Li, Chuangxin Chu, Yujia Zheng, Bohan Zhang et al.ICLR 2026
- MASLeak: Investigating and Exposing Intellectual Property Leakage Vulnerabilities in Multi-Agent SystemsLiwen Wang, Wenxuan Wang, Shuai Wang, Zongjie Li et al.USENIX Security 2026
Builds on12
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
Related papers
- Conjunctive Prompt Attacks in Multi-Agent LLM SystemsNokimul Hasan Arif, Qian Lou, Mengxin ZhengACL 2026
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
- The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMsSergey Berezin, Reza Farahbakhsh, Noël CrespiACL 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- ResMAS: Resilience Optimization in LLM-based Multi-agent SystemsZhilun Zhou, Zihan Liu, Jiahe Liu, Qingyu Shao et al.AAAI 2026 · 2 citations
