Optimizing Agent Planning for Security and Autonomy
Aashish Kolluri, Rishi Sharma, Manuel Costa, Boris Köpf, Tobias Nießen, Mark Russinovich, Shruti Tople, Santiago Zanella-Béguelin
Abstract
Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe actions by enforcing confidentiality and integrity policies, but currently appear costly: they reduce task completion rates and increase token usage compared to probabilistic defenses. We argue that existing evaluations miss a key benefit of system-level defenses: reduced reliance on human oversight. We introduce autonomy metrics to quantify this benefit: the fraction of consequential actions an agent can execute without human-in-the-loop (HITL) approval while preserving security. To increase autonomy, we design a security-aware agent that (i) introduces richer HITL interactions, and (ii) explicitly plans for both task progress and policy compliance. We implement this agent design atop an existing information-flow control defense against prompt injection and evaluate it on the AgentDojo and WASP benchmarks. Experiments show that this approach yields higher autonomy without sacrificing utility. Introduction AI agents are increasingly used in applications ranging from information retrieval (Anthropic, 2025; OpenAI, 2025b; Perplexity, 2025b) to browser and computer-use (OpenAI, 2025a; Perplexity, 2025a; OpenAI, 2025c). These agents often fetch information from various data sources in order to complete user tasks effectively. However, this reliance on external data sources exposes agents to indirect prompt injection attacks (PIAs) (Greshake et al., 2023; Yi et al., 2025), where malicious actors manipulate data sources to hijack the agents' behavior. The security implications of PIAs are particularly critical in scenarios where AI agents are trusted with handling sensitive information, and can manifest e.g. as publishing malicious patches to software packages or the exfiltration of confidential information. Several probabilistic defenses have been proposed against PIAs, such as model alignment (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4e36b0b-3f06-4dc6-a5a2-773932a152f8Builds on8
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff et al.USENIX Security 2026 · 134 citations
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri et al.CCS 2025 · 1 citation
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
Related papers
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou et al.EMNLP 2025
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI AgentsKaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo et al.ICML 2025
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM AgentsFeiran Jia, Tong Wu, Xin Qin, Anna Cinzia SquicciariniACL 2025
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao et al.USENIX Security 2026 · 45 citations
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal AttributionMinbeom Kim, Mihir Parmar, Phillip Wallis, Lesly Miculicich et al.ICML 2026
