CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Yuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, Jet Geronimo, Avi Dhir
Abstract
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture-the-Flag competitions or lack comprehensive coverage. Building a benchmark for realworld vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable attacks. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our experiments show that the state-of-the-art agent framework can exploit up to 13% of the vulnerabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20050955-a2b4-4ec9-9dfd-9288ef9c81f7Cited by top-tier papers15
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai et al.ICLR 2026 · 83 citations
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen et al.ICLR 2026 · 31 citations
- Reliable Weak-to-Strong Monitoring of LLM AgentsNeil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich et al.ICLR 2026 · 28 citations
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration TestingJustin W. Lin, Eliot Jones, Donovan Jasper, Ethan Ho et al.ICLR 2026 · 18 citations
Builds on8
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Data Breaches, Phishing, or Malware?: Understanding the Risks of Stolen CredentialsKurt Thomas, Frank Li, Ali Zand, Jacob Barrett et al.CCS 2017 · 248 citations
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen et al.NeurIPS 2024 · 195 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu et al.USENIX Security 2018 · 138 citations
Related papers
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr et al.ICML 2025
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu et al.ICLR 2026 · 6 citations
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application VulnerabilitiesXiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang et al.ICLR 2026 · 3 citations
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis et al.ICML 2026 · 9 citations
