AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
Haitao Hu, Peng Chen, Yanpeng Zhao, Yuqi Chen
Abstract
Large Language Models (LLMs) have been increasingly integrated into computer-use agents, which can autonomously operate tools on a user's computer to accomplish complex tasks. However, due to the inherently unstable and unpredictable nature of LLM outputs, they may issue unintended tool commands or incorrect inputs, leading to potentially harmful operations. Unlike traditional security risks stemming from insecure user prompts, tool execution results from LLM-driven decisions introduce new and unique security challenges. These vulnerabilities span across all components of a computer-use agent. To mitigate these risks, we propose AgentSentinel, an end-to-end, real-time defense framework designed to mitigate potential security threats on a user's computer. AgentSentinel intercepts all sensitive operations within agent-related services and halts execution until a comprehensive security audit is completed. Our security auditing mechanism introduces a novel inspection process that correlates the current task context with system traces generated during task execution. To thoroughly evaluate AgentSentinel, we present BadComputerUse, a benchmark consisting of 60 diverse attack scenarios across six attack categories. The benchmark demonstrates a 87% average attack success rate on four state-of-the-art LLMs. Our evaluation shows that AgentSentinel achieves an average defense success rate of 79.6%, significantly outperforming all baseline defenses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3304dc78-74ea-469a-a51f-87c507941f55Cited by top-tier papers5
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought CorrectionChangyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong et al.ICML 2026 · 13 citations
- Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent SystemsRishi D. Jha, Harold Triedman, Justin Wagle, Vitaly ShmatikovICLR 2026 · 13 citations
- When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use AgentsYuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye et al.ICML 2026 · 7 citations
- Site Isolation is Dead: How Site Isolation is Broken in Agentic Browsers and ExtensionsSuyoung Lee, Seongho Keum, Changoo Lee, Dongwon Shin et al.S&P 2026
- Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable ReasoningYuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang et al.USENIX Security 2026
Builds on11
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 99 citations
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang et al.USENIX Security 2024 · 83 citations
Related papers
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based AgentsHanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao et al.ICLR 2025
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 33 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety DetectionWeidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee et al.ACL 2025 · 41 citations
- SoK: Attack and Defense Landscape of Agentic AI SystemsJuhee Kim, Wenbo Guo, Dawn SongUSENIX Security 2026
