CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Yuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, Jet Geronimo, Avi Dhir
摘要
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture-the-Flag competitions or lack comprehensive coverage. Building a benchmark for realworld vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable attacks. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our experiments show that the state-of-the-art agent framework can exploit up to 13% of the vulnerabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai 等ICLR 2026 · 被引用 83 次
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen 等ICLR 2026 · 被引用 31 次
- Reliable Weak-to-Strong Monitoring of LLM AgentsNeil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich 等ICLR 2026 · 被引用 28 次
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration TestingJustin W. Lin, Eliot Jones, Donovan Jasper, Ethan Ho 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper8
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Data Breaches, Phishing, or Malware?: Understanding the Risks of Stolen CredentialsKurt Thomas, Frank Li, Ali Zand, Jacob Barrett 等CCS 2017 · 被引用 248 次
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等NeurIPS 2024 · 被引用 195 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu 等USENIX Security 2018 · 被引用 138 次
相关 Paper
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr 等ICML 2025
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 等ICLR 2026 · 被引用 6 次
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application VulnerabilitiesXiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang 等ICLR 2026 · 被引用 3 次
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji 等ICLR 2025
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis 等ICML 2026 · 被引用 9 次
