Cyber-Zero: Training Cybersecurity Agents without Runtime
Terry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar, Zijian Wang
摘要
Large Language Models (LLMs) have achieved remarkable success in software engineering tasks when trained with executable runtime environments, particularly in resolving GitHub issues. However, such runtime environments are often unavailable in other domains, especially cybersecurity, where challenge configurations and execution contexts are ephemeral or restricted. We present CYBER-ZERO, the first runtime-free framework for synthesizing high-quality agent trajectories to train cybersecurity LLMs. CYBER-ZERO leverages publicly available CTF writeups and employs persona-driven LLM simulation to reverse-engineer runtime behaviors and generate realistic, long-horizon interaction sequences without actual environments. Using trajectories synthesized by CYBER-ZERO, we train LLMbased agents that achieve up to 13.1% absolute performance gains over baseline models on three prominent CTF benchmarks: InterCode-CTF, NYU CTF Bench, and Cybench. Our best model, CYBER-ZERO-32B, establishes new state-of-the-art performance among open-weight models, matching the capabilities of proprietary systems like DeepSeek-V3-0324 and Claude-3.5-Sonnet while offering superior cost-effectiveness, and demonstrating that runtime-free trajectory synthesis can effectively democratize the development of state-of-the-art cybersecurity agents. https://github.com/amazon-science/cyber-zero C l a u d e -3 . 7 -S o n n e t C l a u d e -3 . 5 -S o n n e t D e e p S e e k -V 3 -0 3 2 4 G e m i n i -2 . 5 -F l a s h Q w e n 3 -3 2 B Q w e n 3 -1 4 B Q w e n 3 -8 B 0 20 40 60 80 InterCode-CTF C l a u d e -3 . 7 -S o n n e t C l a u d e -3 . 5 -S o n n e t G e m i n i -2 . 5 -F l a s h D e e p S e e k -V 3 -0 3 2 4 * Work done during an internship at Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Training Language Model Agents to Find Vulnerabilities with CTF-DojoTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar 等ICML 2026 · 被引用 12 次
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application VulnerabilitiesXiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang 等ICLR 2026 · 被引用 3 次
- Watermarking LLM Agent TrajectoriesWenlong Meng, Chen GONG, Terry Yue Zhuo, Fan Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper13
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding 等ICML 2024 · 被引用 246 次
相关 Paper
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji 等ICLR 2025
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai 等ICLR 2026 · 被引用 83 次
- Toward Cybersecurity-Expert Small Language ModelsMatan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sa 等ICML 2026 · 被引用 7 次
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li 等ICML 2025 · 被引用 1 次
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
