AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
Hanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li, Guibin Zhang, Kun Wang, Tongliang Liu, Hanan Salam
Abstract
Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, and get confused by unclear safety or security rules. To overcome this evaluation crisis, we introduce AgentAuditor, a universal, training-free, memory-augmented reasoning framework that empowers LLM evaluators to emulate human expert evaluators. AgentAuditor constructs an experiential memory by having an LLM adaptively extract structured semantic features (e.g., scenario, risk, behavior) and generate associated chain-of-thought reasoning traces for past interactions. A multi-stage, contextaware retrieval-augmented generation process then dynamically retrieves the most relevant reasoning experiences to guide the LLM evaluator's assessment of new cases. Moreover, we develop ASSEBench, the first benchmark designed to check how well LLM-based evaluators can spot both safety risks and security threats. ASSEBench comprises 2293 meticulously annotated interaction records, covering 15 risk types across 29 application scenarios. A key feature of ASSEBench is its nuanced approach to ambiguous risk situations, employing "Strict" and "Lenient" judgment standards. Experiments demonstrate that AgentAuditor not only consistently improves the evaluation performance of LLMs across all benchmarks but also sets a new state-of-the-art in LLM-as-a-judge for agent safety and security, achieving human-level accuracy. Our work is openly accessible at https://github.com/Astarojth/AgentAuditor-ASSEBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2dc4970-b49c-4cc6-ae36-6fce9d21ec09Cited by top-tier papers8
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 · 102 citations
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought CorrectionChangyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong et al.ICML 2026 · 13 citations
- Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat SimulationJiongchi Yu, Xiaofei Xie, Qiang Hu, Yuhan Ma et al.NDSS 2026 · 11 citations
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAradhye Agarwal, Gurdit Singh Siyan, Yash Pandya, Joykirat Singh et al.ICML 2026 · 5 citations
- MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory LearningChangyue Jiang, Jiayi Wang, Xin Wen, Jiarun Dai et al.USENIX Security 2026
Builds on28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
Related papers
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based AgentsHanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao et al.ICLR 2025
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- A-MemGuard: A Proactive Defense Framework For LLM-Based Agent MemoryQianshan Wei, Tengchao Yang, Yaochen Wang, Xinfeng Li et al.ICML 2026
- GuardAgent: Safeguard LLM Agents via Knowledge-Enabled ReasoningZhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong et al.ICML 2025
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 246 citations
