LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse K. Coskun, Gianluca Stringhini
摘要
Large Language Models (LLMs) have been suggested for use in automated vulnerability repair, but benchmarks showing they can consistently identify security-related bugs are lacking. We thus develop SecLLMHolmes, a fully automated evaluation framework that performs the most detailed investigation to date on whether LLMs can reliably identify and reason about security-related bugs. We construct a set of 228 code scenarios and analyze eight of the most capable LLMs across eight different investigative dimensions using our framework. Our evaluation shows LLMs provide non-deterministic responses, incorrect and unfaithful reasoning, and perform poorly in real-world scenarios. Most importantly, our findings reveal significant non-robustness in even the most advanced models like ‘PaLM2’ and ‘GPT-4’: by merely changing function or variable names, or by the addition of library functions in the source code, these models can yield incorrect answers in 26% and 17% of cases, respectively. These findings demonstrate that further LLM advances are needed before LLMs can be used as general purpose security assistants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations CentresRonal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mohan Baruwal Chhetri 等S&P 2026 · 被引用 44 次
- Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with JustificationsWei Ma, Daoyuan Wu, Yuqiang Sun, Tianwen Wang 等ICSE 2025 · 被引用 28 次
- Cyber-Zero: Training Cybersecurity Agents without RuntimeTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar 等ICLR 2026 · 被引用 22 次
- Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input GenerationHaoxin Tu, Seongmin Lee, Yuxian Li, Peng Chen 等S&P 2026 · 被引用 22 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- Hackers vs. Testers: A Comparison of Software Vulnerability Discovery ProcessesDaniel Votipka, Rock Stevens, Elissa M. Redmiles, Jeremy Hu 等S&P 2018 · 被引用 151 次
相关 Paper
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 被引用 1 次
- Examining Zero-Shot Vulnerability Repair with Large Language ModelsHammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri 等S&P 2023
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si 等ICML 2026
- IRIS: LLM-Assisted Static Analysis for Detecting Security VulnerabilitiesZiyang Li, Saikat Dutta, Mayur NaikICLR 2025
- Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language ModelsHao Shen, Ming Hu, Jiaye Li, Xiaofei Xie 等ISSTA 2026
