SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned
Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Michael Pelican
摘要
DARPA's AI Cyber Challenge (AIxCC, 2023(AIxCC, -2025) ) is the largest competition to date for building fully autonomous Cyber Reasoning Systems (CRSs) that leverage recent advances in AI-particularly large language models (LLMs)-to discover and remediate vulnerabilities in real-world open-source software. This paper presents the first systematic analysis of AIxCC. Drawing on design documents, source code, execution traces, and discussions with organizers and all finalist teams, we examine the competition's structure and key design decisions, characterize the architectural approaches of finalist CRSs, and analyze competition results beyond the final scoreboard. Our analysis reveals the factors that truly drove CRS performance, identifies genuine technical advances achieved by teams, and exposes limitations that remain open for future research. We conclude with lessons for organizing future competitions and broader insights toward deploying autonomous CRSs in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Contextualizing Sink Knowledge for Java Vulnerability DiscoveryFabian Fleischer, Cen Zhang, Joonun Jang, Jeongin Cho 等S&P 2026 · 被引用 2 次
- SnakeCharmer: Automatic Fuzzing Harness Generation for Pure and Hybrid Python LibrariesGabriel Sherman, Stefan NagyFSE 2026
它引用的顶会 Paper11
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- NAUTILUS: Fishing for Deep Bugs with GrammarsCornelius Aschermann, Tommaso Frassetto, Thorsten Holz, Patrick Jauernig 等NDSS 2019 · 被引用 291 次
- Ijon: Exploring Deep State Spaces via FuzzingCornelius Aschermann, Sergej Schumilo, Ali Abbasi, Thorsten HolzS&P 2020 · 被引用 146 次
相关 Paper
- Rise of the HaCRS: Augmenting Autonomous Cyber Reasoning Systems with Human AssistanceYan Shoshitaishvili, Michael Weissbacher, Lukas Dresel, Christopher Salls 等CCS 2017 · 被引用 57 次
- Measuring and Augmenting Large Language Models for Solving Capture-the-Flag ChallengesZimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma 等CCS 2025
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 等ICLR 2026 · 被引用 6 次
- Examining Zero-Shot Vulnerability Repair with Large Language ModelsHammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri 等S&P 2023
- AICrypto: Evaluating Cryptography Capabilities of Large Language ModelsYu Wang, Yijian Liu, Liheng Ji, Han Luo 等ICML 2026 · 被引用 3 次
