Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE
Benjamin Steenhoek, Kalpathy Sivaraman, Renata Saldivar Gonzalez, Yevhen Mohylevskyy, Roshanak Zilouchian Moghaddam, Wei Le
摘要
Security vulnerabilities impose significant costs on users and organizations. Detecting and addressing these vulnerabilities early is crucial to avoid exploits and reduce development costs. Recent studies have shown that deep learning models can effectively detect security vulnerabilities. Yet, little research explores how to adapt these models from benchmark tests to practical applications, and whether they can be useful in practice. This paper presents the first empirical study of a vulnerability detection and fix tool with professional software developers on real projects that they own. We implemented DeepVulguard, an IDE-integrated tool based on state-of-the-art detection and fix models, and show that it has promising performance on benchmarks of historic vulnerability data. DeepVulguard scans code for vulnerabilities (including identifying the vulnerability type and vulnerable region of code), suggests fixes, provides natural-language explanations for alerts and fixes, leveraging chat interfaces. We recruited 17 professional software developers at Microsoft, observed their usage of the tool on their code, and conducted interviews to assess the tool's usefulness, speed, trust, relevance, and workflow integration. We also gathered detailed qualitative feedback on users' perceptions and their desired features. Study participants scanned a total of 24 projects, 6.9 k files, and over 1.7 million lines of source code, and generated 170 alerts and 50 fix suggestions. We find that although state-of-the-art AI-powered detection and fix tools show promise, they are not yet practical for real-world use due to a high rate of false positives and non-applicable fixes. User feedback reveals several actionable pain points, ranging from incomplete context to lack of customization for the user's codebase. Additionally, we explore how AI features, including confidence scores, explanations, and chat interaction, can apply to vulnerability detection and fixing. Based on these insights, we offer practical recommendations for evaluating and deploying AI detection and fix models. Our code and data are available at this link: https://doi.org/10.6084/m9.figshare.26367139.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PredicateFix: Repairing Static Analysis Alerts with Bridging PredicatesYuan-An Xiao, Weixuan Wang, Dong Liu, Junwei Zhou 等ICSE 2026
- Software Vulnerability Management in the Era of Artificial Intelligence: An Industry PerspectiveM. Mehdi Kholoosi, Triet Huynh Minh Le, M. Ali BabarICSE 2026
- Repairing LLM Executions for Secure Automatic ProgrammingAli El Husseini, Yacine Izza, Blaise Genest, Abhik RoychoudhuryICSE 2026
它引用的顶会 Paper12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu 等ICSE 2024 · 被引用 264 次
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- Large Language Models for Test-Free Fault LocalizationAidan Z. H. Yang, Claire Le Goues, Ruben Martins, Vincent J. HellendoornICSE 2024 · 被引用 98 次
相关 Paper
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao 等FSE 2023 · 被引用 42 次
- VulAdvisor: Natural Language Suggestion Generation for Software Vulnerability RepairJian Zhang, Chong Wang, Anran Li, Wenhan Wang 等ASE 2024 · 被引用 7 次
- VulChecker: Graph-based Vulnerability Localization in Source CodeYisroel Mirsky, George Macon, Michael D. Brown, Carter Yagemann 等USENIX Security 2023
- An empirical study on program failures of deep learning jobsRu Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu 等ICSE 2020 · 被引用 96 次
- DeepCVA: Automated Commit-level Vulnerability Assessment with Deep Multi-task LearningTriet Huynh Minh Le, David Hin, Roland Croft, Muhammad Ali BabarASE 2021 · 被引用 62 次
