AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
Jingwei Sun, Jianing Zhu, Yuanyi Li, Tongliang Liu, Xia Hu, Bo Han
Abstract
Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows. However, real-world execution environments are far from ideal: pop-ups, resolution changes, and competing applications frequently interfere with agent perception and control. We introduce AgentHijack, a benchmark designed to evaluate the robustness of computer-use agents under common corruptions, where the uncertainties in dynamic environment disrupt the execution flow without direct adversarial intent. Specifically, AgentHijack introduces 9 configurable common corruptions to replicate realistic imperfect scenarios. We evaluate a variety of desktop tasks that utilize MLLM-based agents and discover that even minor instances of corruption can result in substantial performance degradation, which emphasizes the fragility of agents and underscores the necessity of robustness evaluation. Afterward, we propose AgentHijack-Agent, a framework that integrates an action generator with enhanced grounding capabilities and an onlooker responsible for behavior summarization and environment checking. Extensive experiments validate its effectiveness. Our code, environment, baseline models and data are publicly available at: https://AgentHijack.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29551fb0-fa5e-4712-8f3a-5df5d036810cBuilds on9
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang et al.ICLR 2026 · 109 citations
- Attacking Vision-Language Computer Agents via Pop-upsYanzhe Zhang, Tao Yu, Diyi YangACL 2025 · 99 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use AgentsHanyu Lai, Xiao Liu, Yanxiao Zhao, Han Xu et al.ICLR 2026 · 45 citations
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanSaaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang et al.ICLR 2025 · 2 citations
Related papers
- WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop EnvironmentsHaoren Zhao, Tianyi Chen, Zhen WangICML 2026
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 33 citations
- Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental DistractionsXinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan et al.ACL 2025 · 54 citations
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use AgentsHaitao Hu, Peng Chen, Yanpeng Zhao, Yuqi ChenCCS 2025
- ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic EnvironmentsRuicheng Ao, Zeping Min, Tingyu Zhu, Wotao Yin et al.ICLR 2026
