AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, J. Zico Kolter, Matt Fredrikson, Yarin Gal, Xander Davies
摘要
Description: Validation dataset for the TELOS runtime AI governance framework against the AgentHarm benchmark (Gray Swan AI / ICLR 2025). 352 adversarial tasks evaluated with two embedding models to demonstrate architecture-level model agnosticism. Key Results — MiniLM (384-dim, local inference): - 352 tasks validated - 74.1% defense success rate (261 blocked, 91 passed) - 1 boundary violation detected - Average latency: 60ms per governance check - Embedding model: sentence-transformers/all-MiniLM-L6-v2 Key Results — Mistral (1024-dim, API inference): - 352 tasks validated - 100% defense success rate (all harmful tasks blocked) - 239 boundary violations detected - Average latency: 351ms per governance check - Embedding model: mistral-embed Why Two Models: The same TELOS governance architecture — identical fidelity calculations, identical thresholds, identical primacy attractor — produces different precision levels depending on embedding dimensionality. MiniLM's 384-dimensional space is insufficient for placing all harmful content far enough from boundary specifications to trigger detection. Mistral's 1024-dimensional space produces sharper geometric separation, resulting in 239 boundary violations versus 1. This validates that TELOS governance is embedding-model-agnostic: the mathematical framework is constant, the measurement precision scales with the embedding model. Files Included: MiniLM Results: - agentharm_forensic_report.json — Aggregate forensic statistics (MiniLM) - agentharm_trace_20260208_220028.jsonl — Per-task JSONL execution traces with governance event log (MiniLM) - agentharm_forensic_report.md — Human-readable forensic summary (MiniLM) - agentharm_exemplar_results.json — Exemplar embedding results (MiniLM) Mistral Results: - agentharm_forensic_report_mistral.json — Aggregate forensic statistics (Mistral) - agentharm_trace_20260208_223516.jsonl — Per-task JSONL execution traces with governance event log (Mistral) - agentharm_forensic_report_mistral.md — Human-readable forensic summary (Mistral) - agentharm_exemplar_mistral_results.json — Exemplar embedding results (Mistral) Cross-Model Comparison: - embedding_comparison_report.json — Detailed side-by-side comparison of governance decisions across both embedding models (127 KB) Benchmark Source: AgentHarm (Gray Swan AI, ICLR 2025) — 352 adversarial tasks designed to test whether AI agents can be manipulated into performing harmful actions including fraud, cyberattacks, and harassment. Published at the International Conference on Learning Representations, 2025. Validation Status: This dataset demonstrates validated runtime governance performance across two embedding architectures. The MiniLM result (74.1% DSR) represents an honest measurement of governance precision at lower embedding dimensionality. The Mistral result (100% DSR) demonstrates that the same governance framework achieves full coverage with higher-dimensional embeddings. Both results are deterministic and reproducible given the same embedding models and governance configuration. The governance engine implementation is proprietary; forensic output data is published for independent analysis. Validation Date: 2026-02-08
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li 等NeurIPS 2025 · 被引用 98 次
- OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent SafetySanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang 等ICLR 2026 · 被引用 75 次
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS EnvironmentsZeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning 等ICLR 2026 · 被引用 46 次
- Towards a Science of AI Agent ReliabilityStephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu 等ICML 2026 · 被引用 45 次
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan 等ACL 2025 · 被引用 37 次
它引用的顶会 Paper13
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
相关 Paper
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic ImagesBo Zhang, Tzu-Yen Ma, Zichen Tang, Junpeng Ding 等ACL 2026
- TRAP: Targeted Redirecting of Agentic PreferencesHangoo Kang, Jehyeok Yeon, Gagandeep SinghNeurIPS 2025 · 被引用 7 次
- DynaGuard: A Dynamic Guardian Model With User-Defined PoliciesMonte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah 等ICLR 2026 · 被引用 18 次
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis 等ICML 2026 · 被引用 9 次
