EVMbench: Evaluating AI Agents on Smart Contract Security
Justin Wang, Andreas Bigger, Xiaohai Xu, Justin W. Lin, Andy Applebaum, Tejal Patwardhan, Alpin Yukseloglu, Olivia Watkins
Abstract
Smart contracts on public blockchains now manage large amounts of value, and vulnerabilities in these systems can lead to substantial losses. As AI agents become more capable at reading, writing, and running code, it is natural to ask how well they can already navigate this landscape, both in ways that improve security and in ways that might increase risk. We introduce EVMbench, an evaluation that measures the ability of agents to detect, patch, and exploit smart contract vulnerabilities. EVMbench draws on 117 curated vulnerabilities from 40 repositories and, in the most realistic setting, uses programmatic grading based on tests and blockchain state under a local Ethereum execution environment. We evaluate a range of frontier agents and find that they are capable of discovering and exploiting vulnerabilities end-to-end against live blockchain instances. We release code, tasks, and tooling to support continued measurement of these capabilities and future work on security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de1bbde9-fbc2-4824-a13c-d6e4e9f286c0Builds on4
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai et al.ICLR 2026 · 83 citations
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes HeideckeICML 2025
Related papers
- CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity CapabilitiesTianneng Shi, Robin Rheem, Dongwei Jiang, Mona Wang et al.ICML 2026 · 3 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security VulnerabilityXianzhen Luo, Jingyuan Zhang, Shiqi Zhou, JinYang Huang et al.ICML 2026 · 3 citations
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li et al.ICML 2025 · 1 citation
- EVMPatch: Timely and Automated Patching of Ethereum Smart ContractsMichael Rodler, Wenting Li, Ghassan O. Karame, Lucas DaviUSENIX Security 2021 · 103 citations
