CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
Xianzhen Luo, Jingyuan Zhang, Shiqi Zhou, JinYang Huang, Chuan Xiao, Qingfu Zhu, Zhiyuan Ma, YUE XING, Yang Yue, WencongZeng, Wanxiang Che
Abstract
Evaluating and improving the security capabilities of code agents requires high-quality, executable vulnerability tasks. However, existing works rely on costly, unscalable manual reproduction and suffer from outdated data distributions. To address these, we present CVE-Factory, the first multi-agent framework to achieve expertlevel quality in automatically transforming sparse CVE metadata into fully executable agentic tasks. Cross-validation against human expert reproductions shows that CVE-Factory achieves 95% solution correctness and 96% environment fidelity, confirming its expert-level quality. It is also evaluated on the latest realistic vulnerabilities and achieves a 66.2% verified success. This automation enables two downstream contributions. First, we construct LiveCVEBench, a continuously updated benchmark of 190 tasks spanning 14 languages and 153 repositories that captures emerging threats including AI-tooling vulnerabilities. Second, we synthesize over 1,000 executable training environments, the first large-scale scaling of agentic tasks in code security. Finetuned Qwen3-32B improves from 5.3% to 35.8% on LiveCVEBench, surpassing Claude 4.5 Sonnet, with gains generalizing to Terminal Bench (12.5% to 31.3%). We open-source all code, data, and models at https://github.com/ livecvebench/CVE-Factory .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3617d8c-416a-460a-87ce-74fe669ff2f3Builds on12
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen et al.FSE 2022 · 206 citations
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu et al.USENIX Security 2018 · 138 citations
- How Effective Are Neural Networks for Fixing Security VulnerabilitiesYi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier et al.ISSTA 2023 · 86 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin et al.ICSE 2025 · 44 citations
Related papers
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An et al.ACL 2026 · 5 citations
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li et al.ICML 2025 · 1 citation
- CVE-Genie: An LLM-Based Multi-Agent Framework for Reproducing CVEsSaad Ullah, Praneeth Balasubramanian, Wenbo Guo, Amanda Burnett et al.CCS 2026
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?Yibo Peng, James Song, Lei Li, Xinyu Yang et al.ACL 2026 · 2 citations
- EVMbench: Evaluating AI Agents on Smart Contract SecurityJustin Wang, Andreas Bigger, Xiaohai Xu, Justin W. Lin et al.ICML 2026 · 8 citations
