USENIX Security2024Top-tier venue
PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing
Gelei Deng, Yi Liu, Víctor Mayoral Vilches, Peng Liu, Yuekang Li, Yuan Xu, Martin Pinzger, Stefan Rass, Tianwei Zhang, Yang Liu
Abstract
Penetration testing, a crucial industrial practice for ensuring system security, has traditionally resisted automation due to the extensive expertise required by human professionals. Large Language Models (LLMs) have shown significant advancements in various domains, and their emergent abilities suggest their potential to revolutionize industries. In this work, we establish a comprehensive benchmark using real-world penetration testing targets and further use it to explore the capabilities of LLMs in this domain. Our findings reveal that while LLMs demonstrate proficiency in specific sub-tasks within the penetration testing process, such as using testing tools, interpreting outputs, and proposing subsequent actions, they also encounter difficulties maintaining a whole context of the overall testing scenario. Based on these insights, we introduce PENTESTGPT, an LLM-empowered automated penetration testing framework that leverages the abundant domain knowledge inherent in LLMs. PENTESTGPT is meticulously designed with three self-interacting modules, each addressing individual sub-tasks of penetration testing, to mitigate the challenges related to context loss. Our evaluation shows that PENTESTGPT not only outperforms LLMs with a task-completion increase of 228.6% compared to the GPT-3.5 model among the benchmark targets, but also proves effective in tackling real-world penetration testing targets and CTF challenges. Having been open-sourced on GitHub, PENTESTGPT has garnered over 6,500 stars in 12 months and fostered active community engagement, attesting to its value and impact in both the academic and industrial spheres.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6741b86c-70dc-48a1-9930-e216dba60702Cited by top-tier papers24
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations CentresRonal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mohan Baruwal Chhetri et al.S&P 2026 · 44 citations
- Incalmo: an Autonomous Llm-Assisted System for Red Teaming Multi-Host NetworksBrian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain et al.S&P 2026 · 29 citations
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar et al.ICLR 2026 · 22 citations
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- Incident Response Planning Using a Lightweight Large Language Model with Reduced HallucinationKim Hammar, Tansu Alpcan, Emil C. LupuNDSS 2026 · 16 citations
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang et al.FSE 2024 · 12 citations
Related papers
- PwnGPT: Automatic Exploit Generation Based on Large Language ModelsWanzong Peng, Lin Ye, Xuetao Du, Hongli Zhang et al.ACL 2025 · 7 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu et al.ICLR 2026 · 6 citations
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration TestingLanxiao Huang, Daksh Dave, Tyler Cody, Peter A. Beling et al.EMNLP 2025 · 1 citation
- Cloak, Honey, Trap: Proactive Defenses Against LLM AgentsDaniel Ayzenshteyn, Roy Weiss, Yisroel MirskyUSENIX Security 2025
