PoCE: Automated Proof-of-Concept Synthesis using Large Language Models for Robust Validation
Tanusree Das Tithy, Lamia Hasan Rodoshi, Ayman Rafid Azahar, Amlan Abhidarshi, Tabassum Faruk, Fahmid Al Rifat, Faysal Hossain Shezan
摘要
Vulnerability reports play a critical role in software repair, with Proof-of-Concept (PoC) tests serving as one of their most essential components. PoC tests enable software developers to reliably reproduce reported vulnerabilities and subsequently deploy patches. However, generating effective PoCs is costly, expertise-intensive, and increasingly challenging due to the diversity of modern software ecosystems and their complex dependencies. Inadequate or incorrect PoCs can significantly delay patch deployment, thereby increasing the window of exposure to attacks. Prior work on automated PoC generation struggles to produce comprehensive and reliable testing. In this work, we present an automated PoC generation framework, PoCE, capable of generating PoCs across diverse software systems by handling varied input formats and complex execution contexts using large language models. PoCE integrates structured in-context learning, retrieval-augmented generation, and iterative chain-of-thought reasoning to expand an initial successful PoC into multiple validated variants. These variants are executed in controlled environments to confirm success. We evaluate PoCE on thirteen widely used software projects, including TensorFlow, Yasm, Zlib, Liblouis, Cflow, Pytorch, Node.js, TCPDUMP, Fig2dev, Binutils, libsndfile, LibTIFF, and libsixel. Our approach achieves a success rate of 77.7% and generates multiple PoC variants for the most vulnerable cases, uncovering alternative trigger paths and edge conditions. We discover 68 zero-day PoCs and identify 26 previously unknown zero-day vulnerabilities in cross-layer software.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm PackagesDeniz Simsek, Aryaz Eghbali, Michael PradelFSE 2026 · 被引用 4 次
- PAGENT: Program Analysis Guided LLM Agent for Proof-of-Concept GenerationAchintya Desai, Md Shafiuzzaman, Wenbo Guo, Tevfik BultanISSTA 2026
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- SemFuzz: Semantics-based Automatic Generation of Proof-of-Concept ExploitsWei You, Peiyuan Zong, Kai Chen, XiaoFeng Wang 等CCS 2017 · 被引用 148 次
- V2E: Validating Smart Contract Vulnerabilities through Profit-Driven Exploit Generation and ExecutionJingwen Zhang, Yuhong Nan, Kaiwen Ning, Mingxi Ye 等FSE 2026
