PAGENT: Program Analysis Guided LLM Agent for Proof-of-Concept Generation
Achintya Desai, Md Shafiuzzaman, Wenbo Guo, Tevfik Bultan
Abstract
Software developers frequently receive vulnerability reports that require them to reproduce the vulnerability in a reliable manner by generating a proof-of-concept (PoC) input that triggers it. Given the source code for a software project and a specific code location for a potential vulnerability, automatically generating a PoC for the given vulnerability has been a challenging research problem. Symbolic execution and fuzzing techniques require expert guidance and manual steps and face scalability challenges for PoC generation. Although recent advances in LLMs have increased the level of automation and scalability, the success rate of PoC generation with LLMs remains quite low. In this paper, we present a novel approach called Program Analysis Guided proof of concept generation agENT (PAGENT) that is scalable and significantly improves the success rate of LLM-based automated PoC generation compared to prior results. PAGENT integrates lightweight and rule-based static analysis phases for providing static analysis guidance and sanitizer-based profiling and coverage information for providing dynamic analysis guidance with a PoC generation agent. Our experiments demonstrate that the resulting hybrid approach significantly outperforms the prior top-performing agentic approach by 132% for the PoC generation task across 10 open-source projects. PAGENT also discovered 32 post-patch PoCs that trigger the vulnerability in the patched version of the source code, with 2 reproducing the crash in the most recent versions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- Where Does It Go?: Refining Indirect-Call Targets with Multi-Layer Type AnalysisKangjie Lu, Hong HuCCS 2019 · 142 citations
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at ScaleZhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai et al.ICLR 2026 · 83 citations
Related papers
- PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm PackagesDeniz Simsek, Aryaz Eghbali, Michael PradelFSE 2026 · 4 citations
- FirmAgent: Leveraging Fuzzing to Assist LLM Agents with IoT Firmware Vulnerability DiscoveryJiangan Ji, Chao Zhang, Shuitao Gan, Lin Jian et al.NDSS 2026 · 12 citations
- PBFuzz: Agentic Directed Fuzzing for PoV GenerationHaochen Zeng, Andrew Bao, Jiajun Cheng, Chengyu SongCCS 2026
- PATCHAGENT: A Practical Program Repair Agent Mimicking Human ExpertiseZheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu et al.USENIX Security 2025
- PoCE: Automated Proof-of-Concept Synthesis using Large Language Models for Robust ValidationTanusree Das Tithy, Lamia Hasan Rodoshi, Ayman Rafid Azahar, Amlan Abhidarshi et al.ISSTA 2026
