VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
Hao Zhu, Jia Li, Cuiyun Gao, Jiaru Qian, Yihong Dong, Huanyu Liu, Lecheng Wang, Ziliang Wang, Xiaolong Hu, Ge Li
Abstract
Large language models (LLMs) have achieved remarkable progress in code understanding and analysis tasks. However, state-of-the-art LLMs demonstrate limited performance in vulnerability detection tasks, and even state-of-the-art models struggle to distinguish vulnerable code from patched code. We argue that a key reason for this limitation is that LLMs lack an understanding of security specifications-the expectations defined by developers and security teams about how code should behave to remain safe. When the actual behavior of the code differs from these expectations and introduces a security risk, it becomes a potential vulnerability. However, such knowledge is rarely explicit in training data, leaving models unable to reason about the root causes of security flaws. To address this challenge, We propose VulInstruct, a specification-guided approach that systematically extracts reusable security specifications from historical vulnerabilities to instruct the detection of new ones. Specifically, VulInstruct designs two automatic pipelines to construct a specification knowledge base from complementary perspectives: (i) General specifications, extracted from high-quality patches across diverse projects, capturing fundamental safe behaviors accumulated across the open-source ecosystem; and (ii) Domain-specific specifications, context-dependent expectations repeatedly violated in particular repositories or domains that are relevant to the target code under analysis. Before analyzing new code, VulInstruct leverages this specification knowledge base to retrieve relevant past cases and their associated specifications, enabling LLMs to reason about expected safe behaviors rather than relying solely on surface patterns. We evaluate VulInstruct under strict evaluation criteria requiring both correct predictions and valid reasoning. On the PrimeVul dataset, VulInstruct achieves 45.0% F1-score (32.7% improvement) and 37.7% recall (50.8% improvement) compared to the strongest baselines, while uniquely detecting 24.3% of all identified vulnerabilities-2.4× more than any baseline. In pair-wise evaluation distinguishing vulnerable from patched code, VulInstruct also achieves a 32.3% relative improvement over the best baseline. Beyond benchmarks, VulInstruct discovered a previously unknown high-severity vulnerability in production code (later assigned CVE-2025-56538) by recognizing violations of extracted specifications, demonstrating its practical value for real-world vulnerability discovery. All code and supplementary materials are available at https://github.com/zhuhaopku/VulInstruct-temp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dca8542-7378-4ba5-9b33-2cd5bda6b412Builds on10
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin et al.ICSE 2025 · 44 citations
- Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code RepositoriesAlperen Yildiz, Sin G. Teo, Yiling Lou, Yebo Feng et al.ACL 2025 · 32 citations
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability DetectionNiklas Risse, Jing Liu, Marcel BöhmeISSTA 2025 · 8 citations
Related papers
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao et al.ISSTA 2025 · 2 citations
- ProSec: Fortifying Code LLMs with Proactive Security AlignmentXiangzhe Xu, Zian Su, Jinyao Guo, Kaiyuan Zhang et al.ICML 2025
- Vul-R2: A Reasoning LLM for Automated Vulnerability RepairXin-Cheng Wen, Zirui Lin, Yijun Yang, Cuiyun Gao et al.ASE 2025 · 1 citation
- SecCodePRM: A Process Reward Model for Code SecurityWeichen Yu, Ravi Mangal, Yinyi Luo, Kai Hu et al.ICML 2026 · 1 citation
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu et al.ISSTA 2024 · 10 citations
