Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems
Rishi D. Jha, Harold Triedman, Justin Wagle, Vitaly Shmatikov
Abstract
Control-flow hijacking attacks manipulate orchestration mechanisms in multi-agent systems into performing unsafe actions that compromise the system and exfiltrate sensitive information. Recently proposed defenses, such as LlamaFirewall, rely on alignment checks of inter-agent communications to ensure that all agent invocations are "related to" and "likely to further" the original objective.
We start by demonstrating control-flow hijacking attacks that evade these defenses even if alignment checks are performed by advanced LLMs. We argue that the safety and functionality objectives of multi-agent systems fundamentally conflict with each other. This conflict is exacerbated by the brittle definitions of "alignment" and the checkers' incomplete visibility into the execution context.
We then propose, implement, and evaluate ControlValve, a new defense based on the principles of control-flow integrity and least privilege. ControlValve (1) generates permitted control-flow graphs for multi-agent systems, and (2) enforces that all executions comply with these graphs, along with contextual rules (generated in a zero-shot manner) for each agent invocation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e67392ea-58d5-48bf-ab41-e9471e8d7da9Cited by top-tier papers1
Ask how each one uses itBuilds on5
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM AgentsHao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li et al.NeurIPS 2025 · 76 citations
- ACE: A Security Architecture for LLM-Integrated App SystemsEvan Li, Tushin Mallick, Evan Rose, William K. Robertson et al.NDSS 2026 · 57 citations
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri et al.CCS 2025 · 1 citation
- IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic SystemsYuhao Wu, Franziska Roesner, Tadayoshi Kohno, Ning Zhang et al.NDSS 2025
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use AgentsHaitao Hu, Peng Chen, Yanpeng Zhao, Yuqi ChenCCS 2025
Related papers
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao et al.USENIX Security 2026 · 45 citations
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang et al.ICML 2026
- Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MASBingyu Yan, Xiaoming Zhang, Ziyi Zhou, Chaozhuo Li et al.AAAI 2026 · 1 citation
- Conjunctive Prompt Attacks in Multi-Agent LLM SystemsNokimul Hasan Arif, Qian Lou, Mengxin ZhengACL 2026
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM AgentsFeiran Jia, Tong Wu, Xin Qin, Anna Cinzia SquicciariniACL 2025
