AIR: Improving Agent Safety through Incident Response
Zibo Xiao, Jun Sun, Junjie Chen
摘要
Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and recent context, (2) guide the agent to execute containment and recovery actions via its tools, and (3) synthesize guardrail rules during eradication to block similar incidents in future executions. We evaluate AIR on three representative agent types. Results show that AIR achieves detection, remediation, and eradication success rates all exceeding 90%. Extensive experiments further confirm the necessity of AIR's key design components, show the timeliness and moderate overhead of AIR, and demonstrate that LLM-generated rules can approach the effectiveness of developer-authored rules across domains. These results show that incident response is both feasible and essential as a first-class mechanism for improving agent safety.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang 等ICML 2024 · 被引用 436 次
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou 等ICLR 2024 · 被引用 303 次
- LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied AgentsJae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim 等ICLR 2024 · 被引用 49 次
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use AgentsJingyi Yang, Shuai Shao, Dongrui Liu, Jing ShaoNeurIPS 2025 · 被引用 33 次
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing 等ICLR 2026 · 被引用 29 次
相关 Paper
- AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety DetectionWeidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee 等ACL 2025 · 被引用 41 次
- Incident Response Planning Using a Lightweight Large Language Model with Reduced HallucinationKim Hammar, Tansu Alpcan, Emil C. LupuNDSS 2026 · 被引用 16 次
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin 等CHI 2026 · 被引用 2 次
- GuardAgent: Safeguard LLM Agents via Knowledge-Enabled ReasoningZhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong 等ICML 2025
- CAST: A Compiler-Based Framework for Systematically Testing LLM Compositional SafetyLu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An 等ISSTA 2026
