Towards Causal Deep Learning for Vulnerability Detection
Md Mahbubur Rahman, Ira Ceka, Chengzhi Mao, Saikat Chakraborty, Baishakhi Ray, Wei Le
摘要
Deep learning vulnerability detection has shown promising results in recent years. However, an important challenge that still blocks it from being very useful in practice is that the model is not robust under perturbation and it cannot generalize well over the out-ofdistribution (OOD) data, e.g., applying a trained model to unseen projects in real world. We hypothesize that this is because the model learned non-robust features, e.g., variable names, that have spurious correlations with labels. When the perturbed and OOD datasets no longer have the same spurious features, the model prediction fails. To address the challenge, in this paper, we introduced causality into deep learning vulnerability detection. Our approach CausalVul consists of two phases. First, we designed novel perturbations to discover spurious features that the model may use to make predictions. Second, we applied the causal learning algorithms, specifically, do-calculus, on top of existing deep learning models to systematically remove the use of spurious features and thus promote causal based prediction. Our results show that CausalVul consistently improved the model accuracy, robustness and OOD performance for all the state-of-the-art models and datasets we experimented. To the best of our knowledge, this is the first work that introduces do calculus based causal learning to software engineering models and shows it's indeed useful for improving the model accuracy, robustness and generalization. Our replication package is located at https://figshare.com/s/0ffda320dcb96c249ef2 . CCS CONCEPTS • Software and its engineering → Empirical software validation; • Security and privacy → Software security engineering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 被引用 63 次
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability DetectionNiklas Risse, Jing Liu, Marcel BöhmeISSTA 2025 · 被引用 8 次
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)Xu Yang, Shaowei Wang, Jiayuan Zhou, Wenhan ZhuFSE 2025 · 被引用 7 次
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou 等FSE 2025 · 被引用 5 次
它引用的顶会 Paper11
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural NetworksSicong Cao, Xiaobing Sun, Lili Bo, Rongxin Wu 等ICSE 2022 · 被引用 100 次
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 被引用 98 次
- Unicorn: reasoning about configurable system performance through the lens of causalityMd Shahriar Iqbal, Rahul Krishna, Mohammad Ali Javidian, Baishakhi Ray 等EuroSys 2022 · 被引用 60 次
相关 Paper
- A Causal Learning Framework for Enhancing Robustness of Source Code ModelsJunyao Ye, Zhen Li, Xi Tang, Deqing Zou 等FSE 2025
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 被引用 54 次
- Adversarial Robustness Through the Lens of CausalityYonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu 等ICLR 2022 · 被引用 65 次
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection SolutionsSatyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad MedvidovicICSE 2025 · 被引用 1 次
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi 等USENIX Security 2024 · 被引用 13 次
