The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
Feiran Jia, Tong Wu, Xin Qin, Anna Cinzia Squicciarini
Abstract
Large Language Model (LLM) agents are increasingly being deployed as conversational assistants capable of performing complex realworld tasks through tool integration. This enhanced ability to interact with external systems and process various data sources, while powerful, introduces significant security vulnerabilities. In particular, indirect prompt injection attacks pose a critical threat, where malicious instructions embedded within external data sources can manipulate agents to deviate from user intentions. While existing defenses show promise, they struggle to maintain robust security while preserving task functionality. We propose a novel and orthogonal perspective that reframes agent security from preventing harmful actions to ensuring task alignment, requiring every agent action to serve user objectives. Based on this insight, we develop Task Shield, a test-time defense mechanism that systematically verifies whether each instruction and tool call contributes to user-specified goals. Through experiments on the Agent-Dojo benchmark, we demonstrate that Task Shield reduces attack success rates (2.07%) while maintaining high task utility (69.79%) on GPT-4o, significantly outperforming existing defenses in various real-world scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfbfd559-4c71-4bb2-af49-6b8a43c155a0Cited by top-tier papers13
- Optimizing Agent Planning for Security and AutonomyAashish Kolluri, Rishi Sharma, Manuel Costa, Boris Köpf et al.ICLR 2026 · 11 citations
- When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use AgentsYuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye et al.ICML 2026 · 7 citations
- AttnTrace: Contextual Attribution of Prompt Injection and Knowledge CorruptionYanting Wang, Runpeng Geng, Ying Chen, Jinyuan JiaS&P 2026 · 6 citations
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow MemoryYuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming et al.CCS 2026
- CachePrune: Teaching LLMs What Not to Follow via KV-Cache EditingRui Wang, Junda Wu, Yu Xia, Tong Yu et al.ACL 2026
Builds on8
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool EmbeddingsShibo Hao, Tianyang Liu, Zhen Wang, Zhiting HuNeurIPS 2023 · 315 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
Related papers
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou et al.EMNLP 2025
- Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated PlanningShanghao Shi, Xiao Wang, Chaoyu Zhang, Hao Li et al.ICML 2026 · 1 citation
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM AgentsHao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li et al.NeurIPS 2025 · 76 citations
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI AgentsKaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo et al.ICML 2025
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao et al.USENIX Security 2026 · 45 citations
