PwnGPT: Automatic Exploit Generation Based on Large Language Models
Wanzong Peng, Lin Ye, Xuetao Du, Hongli Zhang, Dongyang Zhan, Yunting Zhang, Yicheng Guo, Chen Zhang
摘要
Automatic exploit generation (AEG) refers to the automatic discovery and exploitation of vulnerabilities against unknown targets. Traditional AEG often targets a single type of vulnerability and still relies on templates built from expert experience. To achieve intelligent exploit generation, we establish a comprehensive benchmark using Binary Exploitation (pwn) challenges in Capture the Flag (CTF) competitions and investigate the capabilities of Large Language Models (LLMs) in AEG based on the benchmark. To improve the performance of AEG, we propose PwnGPT, an LLM-based automatic exploit generation framework that automatically solves pwn challenges. The structural design of PwnGPT is divided into three main components: analysis, generation, and verification modules. With the help of a modular approach and structured problem inputs, PwnGPT can solve challenges that LLMs cannot directly solve. We evaluate PwnGPT on our benchmark and analyze the outputs of each module. Experimental results show that our framework is highly autonomous and capable of addressing various challenges. Compared to direct input LLMs, PwnGPT increases the completion rate of exploit on our benchmark from 26.3% to 57.9% with the OpenAI o1-preview model and from 21.1% to 36.8% with the GPT-4o model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm PackagesDeniz Simsek, Aryaz Eghbali, Michael PradelFSE 2026 · 被引用 4 次
- From Assistance to Autonomy: An Empirical Study of AI Use in a Live Capture-the-Flag (CTF) CompetitionTingxuan Tang, Nicolas Janis, Kalyn Asher Montague, Kevin Eykholt 等USENIX Security 2026
- Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM AgentsYanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu 等ACL 2026
它引用的顶会 Paper12
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration TestingGelei Deng, Yi Liu, Víctor Mayoral Vilches, Peng Liu 等USENIX Security 2024 · 被引用 186 次
- NAVEX: Precise and Scalable Exploit Generation for Dynamic Web ApplicationsAbeer Alhuzali, Rigel Gjomemo, Birhanu Eshete, V. N. VenkatakrishnanUSENIX Security 2018 · 被引用 85 次
- Revery: From Proof-of-Concept to ExploitableYan Wang, Chao Zhang, Xiaobo Xiang, Zixuan Zhao 等CCS 2018 · 被引用 85 次
相关 Paper
- Measuring and Augmenting Large Language Models for Solving Capture-the-Flag ChallengesZimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma 等CCS 2025
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li 等ICML 2025 · 被引用 1 次
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 等ICLR 2026 · 被引用 6 次
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr 等ICML 2025
- Cloak, Honey, Trap: Proactive Defenses Against LLM AgentsDaniel Ayzenshteyn, Roy Weiss, Yisroel MirskyUSENIX Security 2025
