RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?
Di Cao, Yong Liao, Xiuwei Shang
Abstract
The latest advancements in large language models (LLMs) have sparked interest in their potential for software vulnerability detection. However, there is currently a lack of research specifically focused on vulnerabilities in the PHP language, and challenges in extracting samples and processing persist, hindering the model's ability to effectively capture the characteristics of specific vulnerabilities. In this paper, we present RealVul, the first LLM-based framework designed for PHP vulnerability detection, addressing these issues. By vulnerability candidate detection methods and employing techniques such as normalization, we can isolate potential vulnerability triggers while streamlining the code and eliminating unnecessary semantic information, enabling the model to better understand and learn from the generated vulnerability samples. We also address the issue of insufficient PHP vulnerability samples by improving data synthesis methods. To evaluate RealVul's performance, we conduct an extensive analysis using five distinct code LLMs on vulnerability data from 180 PHP projects. The results demonstrate a significant improvement in both effectiveness and generalization compared to existing methods, effectively boosting the vulnerability detection capabilities of these models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 621ecc48-6698-4ff7-a4fc-8c414206169aCited by top-tier papers1
Ask how each one uses itBuilds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui et al.EMNLP 2023 · 339 citations
- Impact of Code Language Models on Automated Program RepairNan Jiang, Kevin Liu, Thibaud Lutellier, Lin TanICSE 2023 · 164 citations
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 138 citations
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program RepairYuxiang Wei, Chunqiu Steven Xia, Lingming ZhangFSE 2023 · 111 citations
Related papers
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao et al.ISSTA 2025 · 2 citations
- From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability DetectionJie Lin, David MohaisenNDSS 2025
- LLM-based Vulnerability Discovery through the Lens of Code MetricsFelix Weissberg, Lukas Pirch, Erik Imgrund, Jonas Möller et al.ICSE 2026
- Thought Is All You Need: Smart Contract Vulnerability Detection with Thought-Augmented Large Language ModelChaoyuan Peng, Muhui Jiang, Yajin Zhou, Lei WuFSE 2026
- LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language ModelsAhmed Lekssays, Hamza Mouhcine, Khang Tran, Ting Yu et al.USENIX Security 2025
