CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
Minbeom Kim, Mihir Parmar, Phillip Wallis, Lesly Miculicich, Kyomin Jung, Krishnamurthy Dvijotham, Long T. Le, Tomas Pfister
摘要
AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks. In this attack scenario, malicious commands hidden within untrusted content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the over-defense dilemma : they deploy expensive, always-on sanitization regardless of actual threat, thereby degrading utility and latency even in benign scenarios. We revisit IPI through a causal ablation perspective: a successful injection manifests as a dominance shift where the user request no longer provides decisive support for the agent's privileged action, while a particular untrusted segment, such as a retrieved document or tool output, provides disproportionate attributable influence. Based on this signature, we propose CausalArmor , a selective defense framework that (i) computes lightweight, leave-one-out ablation-based attributions at privileged decision points, and (ii) triggers targeted sanitization only when an untrusted segment dominates the user intent. Additionally, CausalArmor employs retroactive Chain-of-Thought masking to prevent the agent from acting on ``poisoned" reasoning traces. We present a theoretical analysis showing that sanitization based on attribution margins conditionally yields an exponentially small upper bound on the probability of selecting malicious actions. Experiments on AgentDojo and DoomArena demonstrate that CausalArmor matches the security of aggressive defenses while improving explainability and preserving utility and latency of AI agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 被引用 118 次
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM AgentsHao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li 等NeurIPS 2025 · 被引用 76 次
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 被引用 37 次
- PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for FreeHao Li, Xiaogeng Liu, Ning Zhang, Chaowei XiaoACL 2025 · 被引用 33 次
相关 Paper
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao 等USENIX Security 2026 · 被引用 45 次
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI AgentsKaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo 等ICML 2025
- Causal Detection of Multi-Step LLM Agent AttacksViraaji Mothukuri, Reza M. PariziICML 2026 · 被引用 2 次
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou 等EMNLP 2025
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM AgentsFeiran Jia, Tong Wu, Xin Qin, Anna Cinzia SquicciariniACL 2025
