Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability Detection
Niklas Risse, Jing Liu, Marcel Böhme
摘要
According to our survey of machine learning for vulnerability detection (ML4VD), 9 in every 10 papers published in the past !ve years de!ne ML4VD as a function-level binary classi!cation problem: Given a function, does it contain a security !aw? From our experience as security researchers, faced with deciding whether a given function makes the program vulnerable to attacks, we would often !rst want to understand the context in which this function is called. In this paper, we study how often this decision can really be made without further context and study both vulnerable and non-vulnerable functions in the most popular ML4VD datasets. We call a function "vulnerable" if it was involved in a patch of an actual security "aw and con!rmed to cause the program's vulnerability. It is "non-vulnerable" otherwise. We !nd that in almost all cases this decision cannot be made without further context. Vulnerable functions are often vulnerable only because a corresponding vulnerability-inducing calling context exists while non-vulnerable functions would often be vulnerable if a corresponding context existed. But why do ML4VD techniques achieve high scores even though there is demonstrably not enough information in these samples? Spurious correlations: We !nd that high scores can be achieved even when only word counts are available. This shows that these datasets can be exploited to achieve high scores without actually detecting any security vulnerabilities. We conclude that the prevailing problem statement of ML4VD is ill-de!ned and call into question the internal validity of this growing body of work. Constructively, we call for more e#ective benchmarking methodologies to evaluate the true capabilities of ML4VD, propose alternative problem statements, and examine broader implications for the evaluation of machine learning and programming analysis research. CCS Concepts: • Security and privacy → Software and application security; • Software and its engineering → Software testing and debugging; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- BlockScan: Detecting Anomalies in Blockchain TransactionsJiahao Yu, Xian Wu, Hao Liu, Wenbo Guo 等NeurIPS 2025 · 被引用 7 次
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An 等ACL 2026 · 被引用 5 次
- Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?Yikun Li, Ngoc Tan Bui, Ting Zhang, Chengran Yang 等ICSE 2026 · 被引用 2 次
- SymRadar: PoC-Centered Bounded Verification for Vulnerability RepairSeungheon Han, YoungJae Kim, Yeseung Lee, Jooyong YiICSE 2026 · 被引用 1 次
它引用的顶会 Paper19
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 被引用 436 次
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 被引用 208 次
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce 等S&P 2024 · 被引用 167 次
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
相关 Paper
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 被引用 63 次
- On the Effectiveness of Function-Level Vulnerability Detectors for Inter-Procedural VulnerabilitiesZhen Li, Ning Wang, Deqing Zou, Yating Li 等ICSE 2024 · 被引用 18 次
- Understanding and Tackling Label Errors in Deep Learning-Based Vulnerability Detection (Experience Paper)Xu Nie, Ningke Li, Kailong Wang, Shangguang Wang 等ISSTA 2023 · 被引用 23 次
- Dos and Don'ts of Machine Learning in Computer SecurityDaniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke 等USENIX Security 2022
- VGX: Large-Scale Sample Generation for Boosting Learning-Based Software Vulnerability AnalysesYu Nong, Richard Fang, Guangbei Yi, Kunsong Zhao 等ICSE 2024 · 被引用 23 次
