VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability Detection
Samiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi, Mona Rahimi
摘要
Despite decades of research in vulnerability detection, vulnerabilities in source code remain a growing problem, and more effective techniques are needed in this domain. To enhance software vulnerability detection, in this paper, we first show that various vulnerability classes in the C programming language share common characteristics, encompassing semantic, contextual, and syntactic properties. We then leverage this knowledge to enhance the learning process of Deep Learning (DL) models for vulnerability detection when only sparse data is available. To achieve this, we extract multiple dimensions of information from the available, albeit limited, data. We then consolidate this information into a unified space, allowing for the identification of similarities among vulnerabilities through nearest-neighbor embeddings. The combination of these steps allows us to improve the effectiveness and efficiency of vulnerability detection using DL models. Evaluation results demonstrate that our approach surpasses existing Stateof-the-art (SOTA) models and exhibits strong performance on unseen data, thereby enhancing generalizability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model ExplanationsMingjie Chen, Tiancheng Zhu, Mingxue Zhang, Yiling He 等FSE 2026
- LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language ModelsAhmed Lekssays, Hamza Mouhcine, Khang Tran, Ting Yu 等USENIX Security 2025
它引用的顶会 Paper13
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
相关 Paper
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao 等FSE 2023 · 被引用 42 次
- Vulnerability Detection with Graph Simplification and Enhanced Graph Representation LearningXin-Cheng Wen, Yupan Chen, Cuiyun Gao, Hongyu Zhang 等ICSE 2023 · 被引用 77 次
- VulChecker: Graph-based Vulnerability Localization in Source CodeYisroel Mirsky, George Macon, Michael D. Brown, Carter Yagemann 等USENIX Security 2023
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- Enhancing Deep Learning-based Vulnerability Detection by Building Behavior Graph ModelBin Yuan, Yifan Lu, Yilin Fang, Yueming Wu 等ICSE 2023 · 被引用 20 次
