LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse K. Coskun, Gianluca Stringhini
Abstract
Large Language Models (LLMs) have been suggested for use in automated vulnerability repair, but benchmarks showing they can consistently identify security-related bugs are lacking. We thus develop SecLLMHolmes, a fully automated evaluation framework that performs the most detailed investigation to date on whether LLMs can reliably identify and reason about security-related bugs. We construct a set of 228 code scenarios and analyze eight of the most capable LLMs across eight different investigative dimensions using our framework. Our evaluation shows LLMs provide non-deterministic responses, incorrect and unfaithful reasoning, and perform poorly in real-world scenarios. Most importantly, our findings reveal significant non-robustness in even the most advanced models like ‘PaLM2’ and ‘GPT-4’: by merely changing function or variable names, or by the addition of library functions in the source code, these models can yield incorrect answers in 26% and 17% of cases, respectively. These findings demonstrate that further LLM advances are needed before LLMs can be used as general purpose security assistants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c6754b8-fa3c-4a04-b20a-9c8055dd8709Cited by top-tier papers35
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations CentresRonal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mohan Baruwal Chhetri et al.S&P 2026 · 44 citations
- Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with JustificationsWei Ma, Daoyuan Wu, Yuqiang Sun, Tianwen Wang et al.ICSE 2025 · 28 citations
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar et al.ICLR 2026 · 22 citations
- Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input GenerationHaoxin Tu, Seongmin Lee, Yuxian Li, Peng Chen et al.S&P 2026 · 22 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
- Hackers vs. Testers: A Comparison of Software Vulnerability Discovery ProcessesDaniel Votipka, Rock Stevens, Elissa M. Redmiles, Jeremy Hu et al.S&P 2018 · 151 citations
Related papers
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 1 citation
- Examining Zero-Shot Vulnerability Repair with Large Language ModelsHammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri et al.S&P 2023
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si et al.ICML 2026
- IRIS: LLM-Assisted Static Analysis for Detecting Security VulnerabilitiesZiyang Li, Saikat Dutta, Mayur NaikICLR 2025
- Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language ModelsHao Shen, Ming Hu, Jiaye Li, Xiaofei Xie et al.ISSTA 2026
