Cyber-Zero: Training Cybersecurity Agents without Runtime
Terry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar, Zijian Wang
Abstract
Large Language Models (LLMs) have achieved remarkable success in software engineering tasks when trained with executable runtime environments, particularly in resolving GitHub issues. However, such runtime environments are often unavailable in other domains, especially cybersecurity, where challenge configurations and execution contexts are ephemeral or restricted. We present CYBER-ZERO, the first runtime-free framework for synthesizing high-quality agent trajectories to train cybersecurity LLMs. CYBER-ZERO leverages publicly available CTF writeups and employs persona-driven LLM simulation to reverse-engineer runtime behaviors and generate realistic, long-horizon interaction sequences without actual environments. Using trajectories synthesized by CYBER-ZERO, we train LLMbased agents that achieve up to 13.1% absolute performance gains over baseline models on three prominent CTF benchmarks: InterCode-CTF, NYU CTF Bench, and Cybench. Our best model, CYBER-ZERO-32B, establishes new state-of-the-art performance among open-weight models, matching the capabilities of proprietary systems like DeepSeek-V3-0324 and Claude-3.5-Sonnet while offering superior cost-effectiveness, and demonstrating that runtime-free trajectory synthesis can effectively democratize the development of state-of-the-art cybersecurity agents. https://github.com/amazon-science/cyber-zero C l a u d e -3 . 7 -S o n n e t C l a u d e -3 . 5 -S o n n e t D e e p S e e k -V 3 -0 3 2 4 G e m i n i -2 . 5 -F l a s h Q w e n 3 -3 2 B Q w e n 3 -1 4 B Q w e n 3 -8 B 0 20 40 60 80 InterCode-CTF C l a u d e -3 . 7 -S o n n e t C l a u d e -3 . 5 -S o n n e t G e m i n i -2 . 5 -F l a s h D e e p S e e k -V 3 -0 3 2 4 * Work done during an internship at Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28d40820-06dc-4b1b-97c4-d4e4d9e9670aCited by top-tier papers3
- Training Language Model Agents to Find Vulnerabilities with CTF-DojoTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar et al.ICML 2026 · 12 citations
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application VulnerabilitiesXiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang et al.ICLR 2026 · 3 citations
- Watermarking LLM Agent TrajectoriesWenlong Meng, Chen GONG, Terry Yue Zhuo, Fan Zhang et al.ICML 2026 · 2 citations
Builds on13
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding et al.ICML 2024 · 246 citations
Related papers
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai et al.ICLR 2026 · 83 citations
- Toward Cybersecurity-Expert Small Language ModelsMatan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa et al.ICML 2026 · 7 citations
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li et al.ICML 2025 · 1 citation
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
