Architecture Matters for Multi-Agent Security
Ben Hagag, William Anderson, Christian Schroeder de Witt, Sarah Scheffler
Abstract
Multi-agent systems (MAS), composed of networks of two or more autonomous AI agents, have become increasingly popular in production deployments, yet introduce security risks that do not arise in single-agent settings. Even if individual agents exhibit robust security, architectural decisions governing their coordination can create attack surfaces that have not been systematically characterized. In this work, we present an empirical study of how MAS design decisions shape the tradeoff between task performance and attack resistance. Across three agentic environments (browser, desktop, and code) and 13 architectural configurations, we use stagewise evaluations that distinguish planning refusal, execution-stage interception, partial harmful execution, and successful attack completion to study three key design choices: (i) agent roles, which determine how authority and responsibility are allocated; (ii) communication topology, which shapes how and when agents interact; and (iii) memory, which determines the context and state visibility accessible to each agent. We find that multi-agent architectures are more vulnerable than standalone agents in the majority of configurations, with attack success rates varying by up to 3.8× at comparable or higher benign accuracy, and that no single design is universally safer. These results motivate the development of further evaluations that move beyond the security properties of a single agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du et al.ICML 2024 · 128 citations
- MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang et al.ACL 2025 · 97 citations
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksMaksym Andriushchenko, Francesco Croce, Nicolas FlammarionICLR 2025 · 7 citations
Related papers
- ACIArena: Toward Unified Evaluation for Agent Cascading InjectionHengyu An, Minxi Li, Jinghuai Zhang, Naen Xu et al.ACL 2026 · 2 citations
- MaMa: A Game-Theoretic Approach for Designing Safe Agentic SystemsJonathan Nöther, Adish Singla, Goran RadanovicICML 2026
- A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent SystemsTianhao Li, Chuangxin Chu, Yujia Zheng, Bohan Zhang et al.ICLR 2026
- Adversarial Attacks On Multi-Agent CommunicationJames Tu, Tsun-Hsuan Wang, Jingkang Wang, Sivabalan Manivasagam et al.ICCV 2021 · 83 citations
- CIA: Inferring the Communication Topology from LLM-based Multi-Agent SystemsYongxuan Wu, Xixun Lin, He Zhang, Nan Sun et al.ACL 2026
