CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song
Abstract
AI agents have significant potential to reshape cybersecurity, making a thorough assessment of their capabilities critical. However, existing evaluations fall short, because they are based on small-scale benchmarks and only measure static outcomes, failing to capture the full, dynamic range of real-world security challenges. To address these limitations, we introduce CyberGym, a large-scale benchmark featuring 1,507 real-world vulnerabilities across 188 software projects. Adjustable to different vulnerability analysis settings, CyberGym primarily tasks agents with generating a proof-of-concept test that reproduces a vulnerability, given only its text description and the corresponding codebase. Our extensive evaluation highlights that CyberGym effectively differentiates agents' and models' cybersecurity capabilities. Even the top-performing combinations only achieve a ∼20% success rate, demonstrating the overall difficulty of CyberGym. Beyond static benchmarking, we show that CyberGym leads to the discovery of 34 zero-day vulnerabilities and 18 historically incomplete patches. These results underscore that CyberGym is not only a robust benchmark for measuring AI's progress in cybersecurity but also a platform for creating direct, real-world security impact. * Indicates equal contribution. 1 CyberGym has been adopted in the system cards of various frontier models for cybersecurity evaluation, such as Claude (Anthropic, a;d;b;e), Kimi (Kimi Team et al., 2026), and GLM (Zeng et al., 2026).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Incalmo: an Autonomous Llm-Assisted System for Red Teaming Multi-Host NetworksBrian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain et al.S&P 2026 · 29 citations
- EVMbench: Evaluating AI Agents on Smart Contract SecurityJustin Wang, Andreas Bigger, Xiaohai Xu, Justin W. Lin et al.ICML 2026 · 8 citations
- PAGENT: Program Analysis Guided LLM Agent for Proof-of-Concept GenerationAchintya Desai, Md Shafiuzzaman, Wenbo Guo, Tevfik BultanISSTA 2026
- BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMsDi Wu, Xu He, Shu Wang, Kun SunUSENIX Security 2026
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao et al.ICML 2026
Builds on15
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 836 citations
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
Related papers
- CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity CapabilitiesTianneng Shi, Robin Rheem, Dongwei Jiang, Mona Wang et al.ICML 2026 · 3 citations
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li et al.ICML 2025 · 1 citation
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar et al.ICLR 2026 · 22 citations
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu et al.ICLR 2026 · 6 citations
