AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, J. Zico Kolter, Matt Fredrikson, Yarin Gal, Xander Davies
Abstract
Description: Validation dataset for the TELOS runtime AI governance framework against the AgentHarm benchmark (Gray Swan AI / ICLR 2025). 352 adversarial tasks evaluated with two embedding models to demonstrate architecture-level model agnosticism. Key Results — MiniLM (384-dim, local inference): - 352 tasks validated - 74.1% defense success rate (261 blocked, 91 passed) - 1 boundary violation detected - Average latency: 60ms per governance check - Embedding model: sentence-transformers/all-MiniLM-L6-v2 Key Results — Mistral (1024-dim, API inference): - 352 tasks validated - 100% defense success rate (all harmful tasks blocked) - 239 boundary violations detected - Average latency: 351ms per governance check - Embedding model: mistral-embed Why Two Models: The same TELOS governance architecture — identical fidelity calculations, identical thresholds, identical primacy attractor — produces different precision levels depending on embedding dimensionality. MiniLM's 384-dimensional space is insufficient for placing all harmful content far enough from boundary specifications to trigger detection. Mistral's 1024-dimensional space produces sharper geometric separation, resulting in 239 boundary violations versus 1. This validates that TELOS governance is embedding-model-agnostic: the mathematical framework is constant, the measurement precision scales with the embedding model. Files Included: MiniLM Results: - agentharm_forensic_report.json — Aggregate forensic statistics (MiniLM) - agentharm_trace_20260208_220028.jsonl — Per-task JSONL execution traces with governance event log (MiniLM) - agentharm_forensic_report.md — Human-readable forensic summary (MiniLM) - agentharm_exemplar_results.json — Exemplar embedding results (MiniLM) Mistral Results: - agentharm_forensic_report_mistral.json — Aggregate forensic statistics (Mistral) - agentharm_trace_20260208_223516.jsonl — Per-task JSONL execution traces with governance event log (Mistral) - agentharm_forensic_report_mistral.md — Human-readable forensic summary (Mistral) - agentharm_exemplar_mistral_results.json — Exemplar embedding results (Mistral) Cross-Model Comparison: - embedding_comparison_report.json — Detailed side-by-side comparison of governance decisions across both embedding models (127 KB) Benchmark Source: AgentHarm (Gray Swan AI, ICLR 2025) — 352 adversarial tasks designed to test whether AI agents can be manipulated into performing harmful actions including fraud, cyberattacks, and harassment. Published at the International Conference on Learning Representations, 2025. Validation Status: This dataset demonstrates validated runtime governance performance across two embedding architectures. The MiniLM result (74.1% DSR) represents an honest measurement of governance precision at lower embedding dimensionality. The Mistral result (100% DSR) demonstrates that the same governance framework achieves full coverage with higher-dimensional embeddings. Both results are deterministic and reproducible given the same embedding models and governance configuration. The governance engine implementation is proprietary; forensic output data is published for independent analysis. Validation Date: 2026-02-08
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aa954ac-0445-460c-b83a-ef35d0193227Cited by top-tier papers47
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li et al.NeurIPS 2025 · 98 citations
- OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent SafetySanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang et al.ICLR 2026 · 75 citations
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS EnvironmentsZeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning et al.ICLR 2026 · 46 citations
- Towards a Science of AI Agent ReliabilityStephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu et al.ICML 2026 · 45 citations
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
Builds on13
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas et al.NeurIPS 2024 · 362 citations
Related papers
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic ImagesBo Zhang, Tzu-Yen Ma, Zichen Tang, Junpeng Ding et al.ACL 2026
- TRAP: Targeted Redirecting of Agentic PreferencesHangoo Kang, Jehyeok Yeon, Gagandeep SinghNeurIPS 2025 · 7 citations
- DynaGuard: A Dynamic Guardian Model With User-Defined PoliciesMonte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah et al.ICLR 2026 · 18 citations
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis et al.ICML 2026 · 9 citations
