When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
Yuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye, Weitong Ruan, Junyi Li, Rahul Gupta, Huan Sun
摘要
Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from the user's original intent. Such misaligned actions may arise from external attacks (e.g., indirect prompt injection) or from internal limitations (e.g., erroneous reasoning). They not only expose CUAs to safety risks, but also degrade task efficiency and reliability. This work makes the first effort to define and study misaligned action detection in CUAs, with comprehensive coverage of both externally induced and internally arising misaligned actions. We further identify three common categories in real-world CUA deployment and construct MisActBench, a benchmark of realistic trajectories with human-annotated, action-level alignment labels. Moreover, we propose DeAction, a practical and universal guardrail that detects misaligned actions before execution and iteratively corrects them through structured feedback. DeAction outperforms all baselines across offline and online evaluations with moderate latency overhead: (1) On MisActBench, it outperforms baselines by over 15% absolute in F1 score; (2) In online evaluation, it reduces attack success rate by over 90% under adversarial settings while preserving or even improving task success rate in benign environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsJaylen Jones, Zhehao Zhang, Yuting Ning, Eric Fosler-Lussier 等ICML 2026 · 被引用 9 次
- Preemptive Detection and Correction of Misaligned Actions in LLM AgentsHaishuo Fang, Xiaodan Zhu, Iryna GurevychEMNLP 2025 · 被引用 1 次
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use AgentsHaitao Hu, Peng Chen, Yanpeng Zhao, Yuqi ChenCCS 2025
相关 Paper
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-DirectednessErfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh 等ICLR 2026 · 被引用 9 次
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use AgentsTri Cao, Bennett Lim, Yue Liu, Yuan Sui 等ICLR 2026 · 被引用 45 次
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM AgentsFeiran Jia, Tong Wu, Xin Qin, Anna Cinzia SquicciariniACL 2025
- MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning CorrectionWenqi Zhang, Yulin Shen, Changyue Jiang, Jiarun Dai 等CCS 2026 · 被引用 3 次
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 被引用 33 次
