Automated unearthing of dangerous issue reports
Shengyi Pan, Jiayuan Zhou, Filipe Roseiro Côgo, Xin Xia, Lingfeng Bao, Xing Hu, Shanping Li, Ahmed E. Hassan
Abstract
The coordinated vulnerability disclosure (CVD) process is commonly adopted for open source software (OSS) vulnerability management, which suggests to privately report the discovered vulnerabilities and keep relevant information secret until the official disclosure. However, in practice, due to various reasons (e.g., lacking security domain expertise or the sense of security management), many vulnerabilities are first reported via public issue reports (IRs) before its official disclosure. Such IRs are dangerous IRs, since attackers can take advantages of the leaked vulnerability information to launch zero-day attacks. It is crucial to identify such dangerous IRs at an early stage, such that OSS users can start the vulnerability remediation process earlier and OSS maintainers can timely manage the dangerous IRs. In this paper, we propose and evaluate a deep learning based approach, namely MemVul, to automatically identify dangerous IRs at the time they are reported. MemVul augments the neural networks with a memory component, which stores the external vulnerability knowledge from Common Weakness Enumeration (CWE). We rely on publicly accessible CVE-referred IRs (CIRs) to operationalize the concept of dangerous IR. We mine 3,937 CIRs distributed across 1,390 OSS projects hosted on GitHub. Evaluated under a practical scenario of high data imbalance, MemVul achieves the best trade-off between precision and recall among all baselines. In particular, the F1-score of MemVul (i.e., 0.49) improves the best performing baseline by 44%. For IRs that are predicted as CIRs but not reported to CVE, we conduct a user study to investigate their usefulness to OSS stakeholders. We observe that 82% (41 out of 50) of these IRs are security-related and 28 of them are suggested by security experts to be publicly disclosed, indicating MemVul is capable of identifying undisclosed dangerous IRs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7eefc9b0-cc99-4a32-8889-9f4a699669c1Cited by top-tier papers10
- Fine-grained Commit-level Vulnerability Type Prediction by CWE Tree StructureShengyi Pan, Lingfeng Bao, Xin Xia, David Lo et al.ICSE 2023 · 30 citations
- Towards More Practical Automation of Vulnerability AssessmentShengyi Pan, Lingfeng Bao, Jiayuan Zhou, Xing Hu et al.ICSE 2024 · 8 citations
- PyRadar: Towards Automatically Retrieving and Validating Source Code Repository Information for PyPI PackagesKai Gao, Weiwei Xu, Wenhao Yang, Minghui ZhouFSE 2024 · 7 citations
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou et al.FSE 2025 · 5 citations
- SCPatcher: Mining Crowd Security Discussions to Enrich Secure Coding PracticesZiyou Jiang, Lin Shi, Guowei Yang, Qing WangASE 2023 · 2 citations
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 273 citations
- Misbehaviour prediction for autonomous driving systemsAndrea Stocco, Michael Weiss, Marco Calzana, Paolo TonellaICSE 2020 · 138 citations
- Identifying Open-Source License Violation and 1-day Security Risk at Large ScaleRuian Duan, Ashish Bijlani, Meng Xu, Taesoo Kim et al.CCS 2017 · 126 citations
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
Related papers
- Silent Taint-Style Vulnerability Fixes IdentificationZhongzhen Wen, Jiayuan Zhou, Minxue Pan, Shaohua Wang et al.ISSTA 2024 · 2 citations
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia et al.ASE 2021 · 84 citations
- SemFuzz: Semantics-based Automatic Generation of Proof-of-Concept ExploitsWei You, Peiyuan Zong, Kai Chen, XiaoFeng Wang et al.CCS 2017 · 148 citations
- VulChecker: Graph-based Vulnerability Localization in Source CodeYisroel Mirsky, George Macon, Michael D. Brown, Carter Yagemann et al.USENIX Security 2023
- Towards the Detection of Inconsistencies in Public Security Vulnerability ReportsYing Dong, Wenbo Guo, Yueqi Chen, Xinyu Xing et al.USENIX Security 2019 · 149 citations
